Google Develops Watermarking for AI-Designed Proteins
Summary
Google DeepMind has published a protein-watermarking system, SynthIDBio, intended to help biosecurity screening distinguish AI-designed proteins from unknown designs. The work addresses a gap in current DNA-order screening: novel proteins generated by AI may not resemble known threats closely enough for existing sequence-based checks to assess them. SynthIDBio adapts Google’s SynthID approach to ProteinMPNN, a widely used AI protein-design tool. As ProteinMPNN selects amino acids along a specified protein backbone, SynthIDBio uses a secret key and the previously selected residues to suggest watermark-compatible amino acids; ProteinMPNN accepts those suggestions only when they remain compatible with a functional protein. The watermark is distributed across the sequence rather than stored in one location, and detection requires the key and a statistical scan of the complete sequence. In tests, the team used the system to design proteins that bind selected natural target proteins, and the watermarked versions retained that binding activity. The result is not a full demonstration of catalytic or other complex functions, but it suggests that watermarking need not automatically disable the designed protein. Google envisions DNA-synthesis providers receiving keys from trusted organizations such as universities or biotechnology companies, allowing them to identify designs from those sources and focus additional review on unrecognized or untrusted AI-designed proteins. The system would simplify screening rather than guarantee that DNA orders are safe. Its effectiveness depends on secure key distribution and maintenance, and short proteins may contain too little watermark signal to detect reliably. A sequence can potentially also be diluted by attaching it to an unwatermarked natural protein. Compatibility is another constraint: ProteinMPNN supports the method, but other AI protein-design systems may use different workflows. Because detection is statistical, the chosen threshold affects false positives and false negatives. The research therefore presents a possible biosecurity tool, while leaving its practical coverage and security properties unresolved.