About
I’m a published author and a former digital court reporter. Four years of producing transcripts that had to survive being checked is where I learned that a record is only worth what it withstands.
I came to evaluation sideways. I had a house editing standard for stripping machine tics out of prose, wanted to know whether it actually worked, and built a harness to find out. The answer was partly no, which turned out to be the interesting part.
prose-eval tests whether a hand-written editorial rubric can separate machine-written prose from human. Structural cadence carries nearly all the signal; the fitted model transfers below chance to a new domain.
bluepencil is the tool that harness justifies. Eighteen gates, every threshold calibrated against 2,602 human documents so the false-positive rate is a measured 5% rather than a guess.
vintage-study asks whether detectors degrade as generators get newer. They don’t. Instruction tuning predicts detectability; release date doesn’t.
prose-cadence-stats publishes the measurements underneath all of it, so anyone can recompute a threshold or challenge one.
Reachable through GitHub.