Skip to content

Build a Human Review Queue from Classification Results

If your output has a suggested label and a score already normalized to 0–1, use this page to practice deciding what should reach a person first. A missing suggested label, or a missing, out-of-range, or below-threshold score goes to review; the remaining suggestions still need spot checks. This tool does not calibrate scores or call them accuracy.

Local workbench / no account needed

Files are read in this browser only. Nothing is uploaded to JevLog or sent to a model. Refreshing clears the current result.

Up to 5 MB; quoted commas and line breaks are supported.

Choose a file or load the demo sample.

The spreadsheet CSV adds a prefix to formula-like cells and may change report values. Download the JSON to process exact report values in another program.

The input CSV uses id, label and confidence. At threshold 0.8, review IDs 2 (0.60), 3 (missing score) and 5 (missing label); keep IDs 1 and 4 for spot checks. Compare the expected full JSON and expected queue CSV. Verify 5 = 3 + 2.

At threshold 0.6, ID 2 equals the threshold and remains for spot checks; the queue contains two records. Missing scores and labels still need review. These are deterministic teaching results, not model-performance evidence.

Are failed requests automatically detected?

Section titled “Are failed requests automatically detected?”

The tool does not inspect status/error fields. A real adapter needs a separate failure queue so stale labels or scores cannot pass into spot checks.

The current generator does not offer online review storage or team assignment. The queue preserves input columns and adds jevlog_review_reason. Store final labels, reviewer identity, reasons and time separately, as shown in the handoff tutorial.

No. Download the full report and queue first, then reload the input and threshold to reproduce them. No login is needed; files are not uploaded and no external actions are executed.

The queue export preserves every original input column and appends a review reason. The on-page table shows the ID, suggested label, score, reason, and a text, feedback, or message column if present. Record the human’s final label, reason, time, and identity separately. Do not overwrite the initial suggestion. This output does not message customers, approve refunds, or execute external actions.

The fictional sample uses 0.8 only as an illustration: three of five records need review (one low score, one missing score, and one missing label), while two remain for spot checks. Choose a real threshold with your own human-reviewed data, then reassess it when rules or the input population change.

Continue with the human review guide, run comparison tool, or CSV preflight checker.