Skip to content
Source companionPublic source reviewed

Build a Source-Linked Research Reading List

A small, searchable table with a topic, review state and source for each record. You have not reproduced Hassan’s project, validated its reported numbers or evaluated scientific quality. Preserve the original references when sharing your own list.

Use caseTurn CSV, documents, or search results into reviewable data

Source · Hassan · @nutlopeSome code6 min

Learn from Hassan’s research-classification example, then try a small, source-linked exercise.

Source · Hassan · @nutlope · X companion

2026-09-21 · Editorial update; no live model call

You have titles, abstracts and notes scattered across saved links. The useful task is not to ask a model to read everything without limits. First create a small table whose records point back to their sources, then assign a manageable set of topics. This guide uses a fictional reading list, not the author’s original dataset.

A source-linked reading workflow Titles / abstracts → IDs + URLs → Topic criteria → Review overlap → Reading list

Classify available evidence; do not invent it.

  • Five to twenty abstracts or notes you are allowed to process.

  • A title or summary column and a stable identifier.

  • Source links retained alongside text, so you can check an unexpected result.

  • Official docs

Start with the included reading_list.csv. Replace it later with your own allowed material. Keep full-text files separate. A missing abstract should be marked as missing rather than silently completed by a model.

2. Write topics before running the classifier

Section titled “2. Write topics before running the classifier”

Define a few useful categories with specific boundaries. Avoid topics so broad that every paper fits. Decide how to handle overlap; for this single-label exercise, choose the primary topic and leave ambiguous records for review.

3. Separate summary writing from classification

Section titled “3. Separate summary writing from classification”

Jev is documented as a structured-decision model. If your workflow needs new prose summaries, treat that as a different step and record its source and cost separately. The bundled Studio demo only classifies the text already present in your file.

Open the Content sample in Studio and choose summary. Run the demo and inspect the record that contains both AI and business terms. Change its label manually. The purpose is to practice a workflow, not reproduce the author’s reported benchmark.

Select a few additional records and label them manually before adjusting your instructions. Keep those records separate from the examples you used to tune the categories. Compare disagreements against the source text; do not treat a plausible category as a verified reading of the paper.

Practice sample: download reading_list.csv (CSV Studio / local practice; not a live run here).

Keep the reading list evidence-linked

Create a working table with a stable ID, title or summary, and a source URL. Retaining the source is more useful than generating an elaborate taxonomy. A record with no abstract should stay visibly incomplete; do not invent an abstract from its title and then classify the invention.

Choose a few categories that serve your reading goal. A paper about AI-assisted sales can overlap technology and business. Decide whether your real workflow supports several labels. This single-label exercise keeps the primary topic and routes ambiguous cases to a person. Relevance to your queue is not scientific quality or factual correctness.

The bundled reading list is fictional and separate from the author’s experiment. Use it to practice field mapping and review, then build a small authorized dataset of your own. Hold back some manually labeled records so that edits to the criteria are not judged only against examples used to tune them.

When sharing a list, keep attribution and links. Publish your own short annotations rather than copying full papers or other people’s summaries. Processing PDFs, generating summaries and assigning topics are separate stages with separate opportunities for errors and costs.

Record Available evidence Action
P01 Agent-recovery abstract AI topic + source
P02 Database indexing Engineering
P03 AI-assisted sales Review overlap
P04 Missing abstract Request information

Fictional records, not the author’s dataset.

A small, searchable table with a topic, review state and source for each record. You have not reproduced Hassan’s project, validated its reported numbers or evaluated scientific quality. Preserve the original references when sharing your own list.

  • Do not publish copyrighted full texts or someone else’s summaries without permission.
  • Do not equate missing metadata with a low-quality paper.
  • The sample categories demonstrate the interface; research-specific criteria need your own evaluation.
  • You can identify the input, its source, and the fields that leave your system.
  • Original identities, failed rows, uncertain cases and human corrections remain visible.
  • You distinguish an offline fixture, an author demo and a live evaluation you ran yourself.

Source boundary: compiled from the linked public sources; not reproduced here. Review classifications before acting; they do not run actions automatically.

This exercise starts with text. PDF extraction needs separate validation.

No. Classification should not be presented as an academic evaluation.