Skip to content

Compare Classification Runs by Stable ID

After changing a rule or model, first learn which records actually changed. Select the ID and label columns for each CSV separately. Header names may differ, but IDs must refer to the same records. The tool aligns records by ID locally, rejects blank or duplicate IDs and uneven rows, and never calls a model.

Local workbench / no account needed

Files are read in this browser only. Nothing is uploaded to JevLog or sent to a model. Refreshing clears the current result.

Up to 5 MB; quoted commas and line breaks are supported.

Choose a file or load the demo sample.

The spreadsheet CSV adds a prefix to formula-like cells and may change report values. Download the JSON to process exact report values in another program.

The previous CSV uses id and label; the new CSV uses record_id and category. Expect 1 changed, 1 missing, 1 added and 1 unchanged record.

ID Previous label New label Difference
2 Technical Other Changed label
3 Other Empty Missing in new run
4 Empty Shipping New in new run

Unchanged ID 1 only contributes to the summary. The expected JSON and downloaded diff contain three differences. Each input has three records, but only two IDs occur in both files; equal sizes do not establish the same comparison scope.

No. Labels use raw-string comparison; differences in case and whitespace count as changes. IDs are trimmed before matching. Keep any label-normalization mapping explicit and versioned.

Why are blank IDs, duplicates or wrong widths rejected?

Section titled “Why are blank IDs, duplicates or wrong widths rejected?”

They prevent a reliable join. Check the CSV structure and original source first. The tool does not pair rows by position. An existing ID with an empty label is different from a missing record.

Does this measure accuracy or store run history?

Section titled “Does this measure accuracy or store run history?”

It only exports differences, without reference-answer metrics or saved history. Download diff-report.csv or JSON before refreshing and separately retain rule versions, model identifiers and input scope.

Result What to check
Changed label Read the source text, old and new rules, and any human reference label.
Missing in new run Check export scope, failures, and filters before blaming the model.
New in new run Confirm the two runs used the same input population before comparing rates.
Unchanged Matching labels still deserve a spot check near the decision boundary.

The fictional sample uses different header names in the two files and produces one changed label, one missing record, and one added record. This is a difference report. Without human-reviewed reference labels it cannot calculate accuracy. Save the rule version, model identifier, input range, and run time alongside the report before deciding whether to ship.

Continue with the full comparison guide, human review queue, or CSV preflight checker.