SynthIDBio Watermarks AI-Generated Proteins Without Disrupting Function
Summary
Researchers at Google DeepMind introduce SynthIDBio, a family of watermarking methods for tracing the provenance of AI-generated protein sequences and structures. SynthIDBio-sequence integrates tournament sampling and watermark-score filtering into ProteinMPNN, which designs amino-acid sequences for generated protein backbones. In laboratory tests, watermarked binders targeting the SARS-CoV-2 receptor-binding domain, VEGF-A and PD-L1 retained binding performance comparable to non-watermarked designs, while a threshold calibrated to a 0.1% false-positive rate achieved 100% true-positive detection in the reported validation set. Across in-silico tests covering 23 targets and 10,000 samples per target, stronger or higher-entropy watermarking improved detectability, but score filtering reduced design pass rates and could increase candidate-generation costs. The sequence watermark can nevertheless be removed effectively by resequencing with ordinary ProteinMPNN; the study estimates lower hit rates after such attacks, with outcomes varying by target and access to structural filters. SynthIDBio-structure fine-tunes the diffusion module and confidence head of AlphaFold 3, training a PointNet-inspired detector alongside the structural objective so the watermark is embedded in model outputs rather than added after sampling. On the AF3 evaluation set, the reported true-positive rate exceeded 99.8% at a 0.1% false-positive rate, while the recommended model showed no reduction in LDDT or template-modelling scores relative to the baseline. Detection remained robust to rigid transformations, coordinate rounding and small noise, but constrained structural relaxation destroyed the watermark. The authors describe SynthIDBio as a zero-bit, technical proof of concept rather than a complete provenance system: it cannot encode detailed identity, sequence watermarking remains vulnerable to resequencing, and both approaches require secret keys, deployment coordination and standards. They discuss possible roles in biosecurity screening and scientific database integrity, while stressing that open tools, false-positive trade-offs and adoption incentives limit immediate operational use.