Vibe Patenting Tests LLM Judges for Professional Patent-Drafting Agents
Summary
The paper introduces Vibe Patenting, an end-to-end testbed for evaluating AI agents that draft patents. A separately invoked LLM judge scores drafts and supplies structured feedback for iterative revision. Across multiple inventions and agent configurations, judge-guided revision consistently improves the judge’s assessed quality, while unguided revision tends to plateau. Feedback also allows a lower-reasoning agent to approach the performance of a substantially more expensive high-reasoning agent. Stronger models, more reasoning, and domain-specific agent workflows each generally improve assessed drafting quality. The authors validate the judge against an independent professional patent attorney and find meaningful agreement, but the strength of that agreement depends heavily on the metric. They also identify systematic calibration differences between the LLM judge and the attorney. The results present LLM judges as potentially useful optimization signals for complex professional workflows, while showing that their measurements require careful interpretation and independent validation.