JevShield Adds a Local Prompt-Injection Firewall for LLM Agents
Summary
JevShield is an open-source, local security layer for LLM applications and agents. It is designed to detect direct and indirect prompt injection, jailbreaks, data-exfiltration attempts, and malicious tool calls before they affect an application. The project uses Ollama's Jev-style decision API with the Nimble decision model to return typed security decisions, probabilities, and risk scores rather than asking a generative model for free-form yes-or-no judgments. A Python policy engine then converts the model output into ALLOW, REVIEW, or BLOCK outcomes, with default thresholds of 0.50 for review and 0.85 for blocking. The system also tracks whether content is trusted or untrusted, marks risky content as tainted, and prevents tainted text from becoming instructions or authorizing tools. JevShield provides FastAPI endpoints, a Streamlit dashboard, simulated tools, an attack corpus, and automated benchmarks covering accuracy, precision, recall, F1, false-positive and false-negative rates, and latency. Its default setup runs Ollama locally and does not send security data to a cloud API. The repository describes the project as a defense-in-depth research and demonstration system: detection is probabilistic, the tools are simulated, and production deployments still require least privilege, authorization, sandboxing, secret isolation, monitoring, and human approval for high-risk actions.