Agent Task Eval
Junior Edda · Updated 4 weeks, 2 days ago
Content
Skill: Agent Task Eval
Capability Domain: TEST_INTEGRATION
Technology Stack: pytest, golden-tasks, optional-deepeval
Linked Activities: TFK-04, TFK-05, TFK-09, DTA-19
Overview
Patterns for lane 4 (task performance quality): frozen tasks with structured ground truth.
Principles
- Structured oracle is source of truth.
- Same harness as lane 1; live provider at temp=0.
- Not a PR merge gate — promotion band vs last certified baseline.
- Deepeval optional overlay for fuzzy SC-01 only.
Golden task catalog
See artifact 56 Part 4.6 — TASK-* under tests/fixtures/agent_tasks/.
Makefile / CI
make test-agent-quality gated on AGENTS_ENABLED; separate agent-eval.yml.
References
Artifact 56 Part 4.6; DTA-19 §16.
Details
- Capability Domain:
- —
- Technology Stack:
- —
- Created:
- 4 weeks, 2 days ago
- Updated:
- 4 weeks, 2 days ago
Playbook
Activities Using This Skill 4
- Define AI Agent Architecture Define Architecture
- Build Fixture Library Test Automation Framework
- Prepare Test Report Test Automation Framework
- Wire CICD Integration Test Automation Framework