Check saved outputs with Promptfoo CSV regression fixtures
A local evaluation archive and a boundary for independent real outputs.
Use caseUse __expected exact assertions and a local JavaScript provider to check saved outputs. An intentional failure verifies reporting without an LLM judge. Includes a complete fixture archive.
Use __expected exact assertions and a local JavaScript provider to check saved outputs. An intentional failure verifies reporting without an LLM judge. Includes a complete fixture archive.
Source · JevLog editorial · Original guide

Source cover · Promptfoo
Published 2026-10-02 from current official docs with original exercises. No live service calls; source publication dates not established.
Problem
Section titled “Problem”Regression checks detect changed behavior. A provider returning expected labels creates fake passes, so store expectations and predictions separately. This archive checks fictional outputs and does not call Jev.
Before you begin
Section titled “Before you begin”-
Install Node.js and Promptfoo and extract promptfoo-local-regression.zip into a new directory.
-
Use equals assertions without an LLM judge. Installation needs network; the provider reads local files only.
Original JevLog diagram, not a product screenshot or measured result.
1. Freeze expected labels
Section titled “1. Freeze expected labels”cases.csv has case_id, text and __expected with equals: labels. c03 deliberately has a wrong prediction to verify reporting. Do not change labels to fit predictions.
case_id,text,__expectedc01,Download crashes,equals:bugc02,Please explain a setting,equals:questionc03,Thanks for your help,equals:other2. Read independent predictions
Section titled “2. Read independent predictions”Read results.json by case ID without reading __expected. Missing IDs return error, not empty classifications. Save as provider.cjs in CommonJS format.
const fs = require('node:fs');const path = require('node:path');module.exports = class SavedResults { id() { return 'saved-results-fixture'; } async callApi(prompt, context) { const results = JSON.parse(fs.readFileSync(path.join(__dirname, 'results.json'), 'utf8')); const id = context?.vars?.case_id; if (!Object.prototype.hasOwnProperty.call(results, id)) return {error: 'Missing result: ' + id}; return {output: results[id], metadata: {mode: 'fictional-local-fixture'}}; }};3. Configure and generate a report
Section titled “3. Configure and generate a report”Keep configuration beside provider.cjs and cases.csv. Run npx promptfoo eval -c promptfooconfig.yaml, then npx promptfoo view. The archive includes these files; this article is not a measured model evaluation.
description: Local saved-output regression fixtureprompts: - '{{text}}'providers: - file://provider.cjstests: file://cases.csv4. Confirm the intentional mismatch
Section titled “4. Confirm the intentional mismatch”results.json predicts question for c03, conflicting with other. If all pass, inspect loaded files and provider logic. Two correct fixtures and one mismatch are not Jev accuracy.
5. Replace fixtures with independently generated outputs
Section titled “5. Replace fixtures with independently generated outputs”Label and scrub your inputs, then generate predictions independently from live calls. Version models, questions and hashes, and retain errors. Missing outputs are failures; measure confidence and review volume separately.
Practice sample: download promptfoo-local-regression.zip for local use; not a live run here.
Result and cautions
Section titled “Result and cautions”A local evaluation archive and a boundary for independent real outputs.
- Old predictions, caching and versions affect comparison; record the exact files.
- Report the intentional c03 mismatch and missing IDs as errors.
- Expectations and predictions are independent, without benchmark claims.
Source boundary: compiled from the linked public sources; not reproduced here. Review classifications before acting; they do not run actions automatically.
Related guides
Section titled “Related guides”- Compare CSV classification results by ID: a worked example
- Choose Jev confidence thresholds with a labeled CSV
- Handle Jev Python timeouts, retries and error states
Does this archive call Jev?
Section titled “Does this archive call Jev?”No. It checks fictional saved outputs. Generate predictions independently and retain errors and versions when integrating.