Essay evaluator
invented rubrics → every score cites its source
LLM-as-judge with retrieval over real admissions rubrics, so a score points at the sentence it came from.
Eight years of Python and Django. Three of us, shipping tools licensed to pharmaceutical partners.
No writing matches those filters.
invented rubrics → every score cites its source
LLM-as-judge with retrieval over real admissions rubrics, so a score points at the sentence it came from.
Buildx, platform flags, and the one emulation cost that turned out to be worth paying.
Why the interview answer and the production answer are not the same answer.
Thresholds that catch real problems without blocking a release nobody can wait on.
Research code and production code fail differently. Both fail at the handoff.
scanning at every stage → median build time unchanged
Dependency, secret and static analysis wired into GitLab CI for a three-person team, with thresholds nobody has needed to switch off.
three days per dataset → one request, answered while you wait
Immunologists ranked epitopes by merging prediction output by hand. Now a public IEDB tool, licensed to pharmaceutical partners.
Showing everything, newest first. Writing →
Backend and platform work. I take research and prototype code the last mile: to a web tool, a CLI, an API, or a container a stranger can install and run.
Research and prototype code that works for the person who wrote it, turned into something a stranger can install and run. That is the work I want.