Senior Software Development Engineer in Test
Owns automated quality engineering for AI Voice Agent services, spanning UI, backend, APIs, and audio/text interactions. The role requires 6+ years of software engineering or SDET experience, strong programming skills, cloud-native testing expertise, and hands-on experience with LLM systems and evaluation.
About the job
Responsibilities
- Own end-to-end quality for agentic features and workflows, including strategy, development, execution, and release qualification.
- Design and build automation tooling and frameworks for AI/LLM-driven systems, including prompt flows, agent orchestration, and tool integrations.
- Develop and maintain evaluation frameworks to measure response quality, accuracy, and hallucination rates.
- Drive automation coverage of 80%+ for critical AI workflows using deterministic and probabilistic validation approaches.
- Integrate AI quality checks into CI/CD pipelines with PR validation feedback cycles under 15 minutes.
- Build tooling for LLM observability and debugging, including prompt tracing and response analysis.
- Partner with Applied AI teams on prompt engineering, model selection, and evaluation strategies.
- Design and execute performance and load tests for AI services, including latency, throughput, and cost efficiency.
- Identify and mitigate risks related to hallucinations, bias, safety, and edge cases.
- Define and track AI quality KPIs such as task success rates, precision, recall, and latency.
- Participate in design and architecture reviews to ensure systems are testable, observable, and resilient.
- Mentor engineers and contribute to AI quality engineering practices.
Requirements
- 6+ years of experience in software engineering or SDET roles, with an emphasis on software development.
- Strong programming skills in Python, Java, or JavaScript.
- Experience testing distributed, cloud-native SaaS systems and APIs.
- Experience coding with AI agents to accelerate development and improve code quality.
- Hands-on exposure to LLMs or AI/ML systems, such as OpenAI, Claude, or Gemini.
- Understanding of non-deterministic systems and probabilistic testing approaches.
- Experience building test frameworks and scalable automation systems.
- Familiarity with AI evaluation techniques, including benchmarking, golden datasets, and human-in-the-loop validation.
- Experience with CI/CD pipelines such as Jenkins or GitHub Actions.
- Strong collaboration skills across distributed teams and time zones.
- Bachelor’s degree in Computer Science or equivalent practical experience.
Technical Stack
- Backend: Python, Go, Google Cloud Platform, Cloud Run, App Engine, Kubernetes, Datastore, Redis, Elasticsearch
- Frontend: Vue 3, React
- AI: LLM APIs, LiveKit, prompt orchestration frameworks, evaluation tooling
Compensation and Benefits
- Base salary range for British Columbia, Canada: $150,500–$175,250 CAD.
- Salary range reflects base salary only and excludes bonus, equity, and benefits.
- Competitive benefits, training, AI tools, and growth opportunities.
Skills
Python, Java, JavaScript, GCP, Cloud Run, App Engine, Kubernetes, Datastore, Redis, Elasticsearch, Vue 3, React, LLM APIs, Livekit, GitHub Actions
Similar jobs
QA Engineering jobsLeads quality engineering for AI Contact Center systems by building scalable automation frameworks and API, UI, capacity, security, and performance tests. Requires 6+ years of software development experience, strong Python or Java skills, and expertise in cloud services, REST testing, and CI integration.
Senior Quality Engineer partnering across product teams to assess risk, build web and API test automation, investigate technical failures, and improve quality practices. Requires strong Playwright and TypeScript experience, exploratory testing skills, CI/CD knowledge, and the ability to debug application code.
Build and own the testing frameworks, integration environments, performance tooling, and resilience capabilities that enable reliable AI inference infrastructure. The role requires strong Go or Python skills, distributed-systems testing experience, Kubernetes and Docker expertise, and the ability to improve test reliability at scale.
Conçoit et maintient des systèmes de test logiciels, électriques et matériels pour valider des vélos et trottinettes en production. Le rôle exige une formation technique, de l’expérience en automatisation, systèmes embarqués, firmware et dépannage directement sur le plancher de fabrication.