Back to News
RSS feedgerbenrijpkema.substack.com

What Control Engineering Teaches Us About Building AI Systems

Summary

Gerben Rijpkema argues that the common goal of building “ChatGPT, but it knows our data” fails when teams treat data connection as a general-purpose solution. Drawing on his work across AI projects, he says successful systems used architectures designed for specific business questions, while many failed prototypes relied on one-size-fits-all tools. His control-engineering analogy starts with the idea that both drones and LLM systems operate under uncertain dynamics and noisy inputs: documents can be incomplete, scattered, or contradictory, and an LLM can produce different outputs for the same input. A simple recipe-book RAG chatbot illustrates the problem. Retrieving the three most similar chunks may answer a request for a recipe, but cannot reliably count every lasagna in the book when only three chunks reach the model; prompting cannot restore missing signal. Claims handling likewise needs contradiction-focused extraction, while drafting insurer replies needs a consistent, curated handbook, so the two tasks require different architectures. Rijpkema proposes a filter stage that extracts and structures the signal needed by a particular use case. Examples include classifying statements in claims, organizing tender documents by apartment, and selectively writing agent memory. Possible implementations include metadata extraction, classification, reranking, deduplication, chunking, and query rewriting. He then recommends closing the loop through agentic retrieval: the system evaluates retrieved information and searches again when necessary. He reports major performance improvements in several projects, while warning that self-assessment is difficult to stabilize. Structural or numeric invariants, such as extracted unit counts matching a project total, can provide stronger feedback and enable self-correction. Too much effort can make an agent iterate indefinitely; too little can make it answer before gathering enough evidence. For broader “company brain” systems, he separates data quality from signal quality, arguing that aggressive curation can erase useful historical or contradictory information. His proposed direction is a catalog connecting scattered sources and recording ownership, relationships, and supersession, combined with dynamic filters selected for each question. The conclusion is that files are measurements, not signal: meaning emerges relative to the question being asked.