QA Automation Engineer – Agentic AI & RAG Systems
Knowledge Artisans Private Limited
Must Have Skills
Preferred Skills
Job Title: QA Automation Engineer – Agentic AI & RAG Systems
Location: Bangalore
Experience: 3–5 Years
Technology Stack: Python, AWS (Step Functions, Bedrock, SageMaker), OpenSearch
Role Overview
We are seeking a highly skilled QA Automation Engineer to drive quality assurance for our Agentic AI orchestration platform. Unlike traditional QA roles, this position focuses on validating non-deterministic AI systems, ensuring the reliability, safety, and accuracy of LLM-driven agents and Retrieval-Augmented Generation (RAG) pipelines.
The ideal candidate will play a critical role in developing testing frameworks for AI reasoning validation, hallucination detection, and workflow orchestration. You will also contribute to building a Human-in-the-Loop (HITL) validation framework using Amazon SageMaker Ground Truth to ensure high-accuracy outputs for financial reporting processes such as ETF AUM calculations.
Key Responsibilities
AI Logic & Response Validation
- Design and implement automated test frameworks to detect LLM hallucinations and evaluate the reasoning quality of AWS Bedrock agents.
- Validate AI-generated responses for accuracy, consistency, and grounding in source data.
RAG System Testing
- Benchmark and validate retrieval accuracy and semantic relevance of data retrieved from Amazon OpenSearch.
- Ensure RAG pipelines correctly combine retrieved data with LLM-generated responses.
Human-in-the-Loop (HITL) Framework
- Configure and manage Amazon SageMaker Ground Truth labeling jobs.
- Define confidence thresholds and workflows that route low-confidence AI outputs to human reviewers for validation.
Workflow & Orchestration Testing
- Validate complex AWS Step Functions workflows, ensuring proper state transitions, exception handling, and system reliability during agent tool calls.
- Test integrations across AWS services such as Lambda, API Gateway, and S3.
AI Safety & Adversarial Testing
- Conduct red-team testing by designing adversarial or “jailbreak” prompts.
- Validate Bedrock Guardrails for protection against PII exposure, restricted content, and unsafe outputs.
Quality Metrics & Reporting
- Develop and maintain quality dashboards tracking:
- Test pass rates
- Hallucination rates
- Retrieval accuracy
- Human intervention frequency
- Provide regular QA reports to engineering and management teams.
Required Skills & Qualifications
- Automation & Testing
- Strong expertise in Python-based automation testing.
- Hands-on experience with Pytest or similar testing frameworks.
- Experience testing APIs and backend systems.
- AWS Ecosystem
- Practical experience with AWS Lambda, S3, API Gateway, IAM policies, and Step Functions.
- Human-in-the-Loop Systems
- Hands-on experience with Amazon SageMaker Ground Truth or other data labeling and verification platforms.
- Vector Search & RAG Concepts
- Basic understanding of vector databases such as OpenSearch.
- Experience validating semantic search and retrieval pipelines.
- Analytical Thinking
- Ability to design and maintain Golden Datasets for regression testing of LLM prompts and outputs.
Preferred Qualifications
- Experience with LLM evaluation frameworks and prompt evaluation methodologies.
- Familiarity with agent orchestration frameworks such as LangGraph or Strands.
- Background in financial services, data extraction pipelines, or ETF-related reporting systems.
Apply to this job
Required — upload a file or paste the text
