Tinfoil describes a plan to combine private AI conversations with automated safety enforcement by running open-weight models and safeguard models inside hardware secure enclaves. The company says conversations remain invisible outside the user’s enclave, while a flagged event exposes only the account-linked conversation ID and not the conversation or violation category. Its proposed rollout uses two layers: models must first stay at or below a 5% failure rate on 297 hard-no prompts covering self-harm encouragement, mass violence and terrorism, and child endangerment; runtime safeguard models then inspect model responses for policy violations. Tinfoil says the safeguards will be open source and verifiable through enclave attestation, allowing auditors to check both the code and the enforced policy. The company acknowledges that it cannot tune safeguards on customer data or conduct human review of flagged chats, so it emphasizes testing and a near-zero false-positive rate. In evaluations, 1,500 prompts from MLCommons AILuminate and HarmBench were narrowed to 297 unambiguous hard-no prompts. The post also reports experiments on 1.2 million WildChat conversations: gpt-oss-safeguard initially flagged 1,402 conversations, while a Kimi-K3 second pass kept 846, implying a 40% false-positive rate for the first pass. On a 50,000-conversation comparison, gpt-oss flagged 578 items and gpt-oss-safeguard flagged 134, with manual review finding the additional gpt-oss flags incorrect. Tinfoil says safeguards run asynchronously and do not add response latency, but flagged chats are stopped, repeated violations can lead to suspension, and no conversation data persists across safeguard-enclave versions. The company also warns that transparent policies can be optimized against and that benchmark performance cannot guarantee safe behavior in all real-world conversations.
