← Back
ATC cover

ATC · Nov 2025 - May 2026 · Madison, WI - Remote

AI Engineer

Owned multi-agent LLM workflows across 7 enterprise use cases spanning knowledge retrieval and compliance. Early accuracy was fine in demos and ugly in logs, so evaluation became the product.

Seven structured evaluation cycles per use case, fine-tuning against the failures each cycle surfaced, and SQL/Python observability over interaction logs so drift showed up before users saw it.

  • 88% inference accuracy, hallucinations held under 4%
  • Autonomous agents across 6 decision paths at 72% end-to-end task completion
  • Responsible-AI guardrails at 66% measured production edge-case coverage
  • Observability pipelines in SQL and Python surfacing drift and failure patterns
  • 7 structured evaluation cycles per use case to keep improving reliability
Multi-agentGuardrailsEvaluation