Observability and Evaluation Engineer
Salary not listed
Mphasis · Charlotte, NC · Full-time
Maintain · Found · posted 29 days ago
ApplyOpens this job on LinkedIn in a new tab, where Mphasis posted it.
Experience asked for: not stated
This posting never puts a number on it, or names several that contradict each other. Read the requirements below before you rule yourself out — we would rather say nothing than guess at it.
An Observability and Evaluation Engineer (often specialized in AI/LLM systems) designs the tracking, tracing, and testing frameworks that monitor production behavior, measure output quality, and catch regressions in complex software or agentic AI workflows. [1, 2, 3, 4]
Core Responsibilities
• Build Observability Pipelines: Instrument applications and LLM/agent pipelines using telemetry tools (like OpenTelemetry, Arize, Galileo, or LangSmith) to capture execution traces, run logs, and latency metrics. [1]
• Design Evaluation Frameworks: Create offline and online evaluation suites to benchmark accuracy, groundedness, toxicity, tool-use correctness, and reasoning-chain validity. []
• Implement Regression & Drift Testing: Build automated test harnesses and continuous evaluation gates to detect model degradation, data drift, or output anomalies before releases reach production. [1]
• Root-Cause Analysis: Investigate execution traces and multi-turn interaction failures to diagnose erratic system behaviors, API misparameters, or bottlenecks. []
• Optimize Performance & Cost: Monitor and balance operational telemetry relating to token consumption, execution speed, and infrastructure costs. [1]
Key Skills & Requirements
• Programming: Strong proficiency in Python or Go for building custom evaluation scripts and hooking into telemetry SDKs.
• Tools & Stacks: Experience with OpenTelemetry, Prometheus, Grafana, or dedicated AI observability platforms (e.g., Arize Phoenix, LangChain/LangSmith, Galileo, Braintrust).
• Testing Methodologies: Designing adversarial prompts, red-teaming protocols, and golden datasets for LLM/system validation.
• System Design: Understanding of distributed microservices or LLM orchestration layers (LangChain, LlamaIndex, AutoGen). [1, 2, 3]
If you are tailoring this for a specific application, would you like me to focus this job description more heavily on traditional cloud/microservices observability or Generative AI / LLM agent evaluation?