api/fixtures at 2b7d302ef28397e797d23ef2fd525d63fa29c2b1 - api

Files

Roberto Musso d3f7099d93 refactor(eval): 3-mode eval harness (step1/step2/full) with Langfuse fixes

- Rewrite eval config with EvalMode (step1, step2, full) replacing prompt_variants
- Rewrite runner with _run_step1, _run_step2, _run_full dispatch
- CLI: replace --variants with --mode flag
- Add 3 fixture YAMLs: classify_invoices (step1), process_invoices (step2), full_invoices (full)
- Remove old freelance_invoices fixture
- Langfuse: mode-aware dataset items (classifications for step1, extraction for step2, both for full)
- Langfuse: link both prompts (batch_file_classifier + batch_processing) in full mode
- Langfuse: post separate classification_precision/recall/f1 scores for full mode
- Langfuse: skip misleading field_accuracy=0 when field_scores is empty (step1)
- Langfuse: include step1_results in trace output
- MockExecutor: mock async_session to bypass DB in full mode
- Journey fixture: remove user_messages (only interactive test kept)

2026-03-24 16:18:51 +01:00

sample_files/invoices

feat(batch-agent): add E2E evaluation harness with Langfuse integration

2026-03-23 08:54:19 +01:00

classify_invoices.yaml

refactor(eval): 3-mode eval harness (step1/step2/full) with Langfuse fixes

2026-03-24 16:18:51 +01:00

full_invoices.yaml

refactor(eval): 3-mode eval harness (step1/step2/full) with Langfuse fixes

2026-03-24 16:18:51 +01:00

journey_invoice_setup.yaml

refactor(eval): 3-mode eval harness (step1/step2/full) with Langfuse fixes

2026-03-24 16:18:51 +01:00

process_invoices.yaml

refactor(eval): 3-mode eval harness (step1/step2/full) with Langfuse fixes

2026-03-24 16:18:51 +01:00