I write, and I build measurement tools for writing. These are the reviews: short, evidence-first pieces about how machine-generated-text detection gets evaluated, and where the evaluation goes wrong.

Every post links to the code and the data behind it. Where a result made my own work look worse, that result is in the post.

Benchmark reviews

  • What GDPval's rubrics actually measure

    GDPval asks whether models can do economically valuable knowledge work. It contains 220 tasks drawn from 44 occupations, each with a prompt, supporting files, a reference deliverable produced by a working professional, and a grading rubric. It is one of the most cited benchmarks in the field, it appears in frontier model release announcements, and its dataset has been downloaded more than 120,000 times.

  • I ran the vintage study. Release date does not predict detectability.

    In the last review I complained that nobody had measured the obvious thing: take a detector, plot its accuracy against the release date of the model that produced the text, and see whether the curve falls. I said it was a week of API calls and a plot.

  • The detection field built its replacements and never switched

    I spent this week building an evaluation harness for a prose rubric, and the corpus made a liar out of it.

subscribe via RSS