DeepMind Builds SynthID Watermarks Into AI-Designed Proteins: The Synthesized Molecules Still Bind Their Targets, and Can Still Be Detected

On September 30 Google DeepMind released SynthID Bio, extending the SynthID watermark it uses for images, audio and text to synthetic biology, with a paper, "Function-preserving watermarking of AI-generated proteins," published in Nature the same day. It subtly guides the choice of amino acids when designing protein sequences and adjusts atomic coordinates in predicted 3D structures, embedding a statistical signal that is invisible to the eye and to routine analysis but detectable by authorized software. The team designed binders with AlphaProteo plus a watermark-enabled ProteinMPNN and had Adaptyv Bio run wet-lab tests on three targets, VEGF-A, the receptor-binding domain of the SARS-CoV-2 spike protein and PD-L1; watermarked designs matched unwatermarked ones on hit rate, binding affinity and sequence diversity, and DeepMind calls them the first watermarked, biologically functional protein binders. The sequence watermark can still be detected after the protein is physically synthesized; the structure watermark is written into model weights by fine-tuning AlphaFold 3's diffusion network and applies only to predicted coordinates. DeepMind also worked with Stanford University and the Arc Institute to add the watermark to bacteriophage genomes designed by Evo 2, and early culture tests showed the phages remained functional. Code, weights and experimental data are available to researchers.

Why proteins need a watermark

Before a DNA synthesis company turns a customer's sequence into a real molecule, it checks the sequence against databases of known pathogens and toxins. That screening rests on an assumption: an unfamiliar sequence most likely comes from a natural organism nobody has studied yet. AI protein design breaks that assumption, because it can produce molecules that resemble no known sequence yet may have similar functions.

SynthID Bio's idea is that instead of trying to prove an unfamiliar sequence is safe, well-behaved models sign their own output. When a screener detects the watermark, it knows the sequence came from a model with safeguards built in; when public databases ingest entries, AI-designed ones can identify themselves rather than being mistaken for natural sequences and distorting later research.

The hard part is not breaking the function

Changing a few pixels in an image does no harm, but changing one amino acid can stop a protein from folding. The paper's most important result is wet-lab evidence on three real targets that watermarked binders bind almost exactly as well as unwatermarked ones.

The two watermarks need to be kept apart. The sequence watermark travels with the molecule and can still be measured after synthesis. The structure watermark exists only in predicted 3D coordinates; a real protein keeps moving in solution and doesn't preserve a fixed set of coordinates, so it can't be detected after synthesis.

What it can't cover

DeepMind itself calls this a proof of concept and one layer of a multi-layer defense, not a solution. It can only mark output from models that choose to use the watermark: someone who switches to an unwatermarked model or redesigns the sequence may remove it, and the absence of a watermark doesn't show that a sequence is natural or safe. Real use in DNA synthesis screening and databases will need synthesis companies, databases and model developers to coordinate on adopting a standard.

Scope note: experimental results come from DeepMind's blog and the Nature paper; the phage experiments are early results; for the licenses of the open-sourced parts, see the official repository.

via: Google DeepMind blog, Google blog, Nature paper, Nature news