27 March 2026

10th, 22nd, 1st

Me alone, my agent system alone, and the two together, on the same leaderboard.

ainlpharness-engineering

Read the paper (PDF)

Imperial’s NLP course runs SemEval-2022 Task 4 as a competition: detect patronising and condescending language aimed at vulnerable people. 340 of us, one leaderboard. Working by hand, without AI support as the department required, I built a DeBERTa-v3 model with a community-aware input prefix, focal loss for a 9.5:1 class imbalance, and a tuned decision threshold. F1 0.59 on the official dev set, against a RoBERTa baseline of 0.48. It came 10th.

Then, to benchmark it, I ran my data science agent system on the same task. Unsupervised, on GPT-4o, one shot, no prompt engineering. Then I put myself back in the loop.

The numbers

  • Me alone: 10/340
  • My system alone: 22/340
  • Me and the system: 1/340

The middle number is the one people skip. Left alone, the agent system was above average and nothing more: 21 entries finished ahead of it, mine among them. If you had shown me only that result I would have concluded agents were not ready for this.

Putting myself back in meant directing the agents, refining the prompts and checking what came out. It also meant building what that needed: guardrails, logging on every agent, and orchestration I could follow, because otherwise I could not tell a reasoned output from an invented one.

I did not have a word for this at the time. The industry settled on one later.

What I would not have guessed is how far the two together finished ahead of either one alone.