Why proteins need a watermark
Before a DNA synthesis company turns a customer's sequence into a real molecule, it checks the sequence against databases of known pathogens and toxins. That screening rests on an assumption: an unfamiliar sequence most likely comes from a natural organism nobody has studied yet. AI protein design breaks that assumption, because it can produce molecules that resemble no known sequence yet may have similar functions.
SynthID Bio's idea is that instead of trying to prove an unfamiliar sequence is safe, well-behaved models sign their own output. When a screener detects the watermark, it knows the sequence came from a model with safeguards built in; when public databases ingest entries, AI-designed ones can identify themselves rather than being mistaken for natural sequences and distorting later research.
The hard part is not breaking the function
Changing a few pixels in an image does no harm, but changing one amino acid can stop a protein from folding. The paper's most important result is wet-lab evidence on three real targets that watermarked binders bind almost exactly as well as unwatermarked ones.
The two watermarks need to be kept apart. The sequence watermark travels with the molecule and can still be measured after synthesis. The structure watermark exists only in predicted 3D coordinates; a real protein keeps moving in solution and doesn't preserve a fixed set of coordinates, so it can't be detected after synthesis.
What it can't cover
DeepMind itself calls this a proof of concept and one layer of a multi-layer defense, not a solution. It can only mark output from models that choose to use the watermark: someone who switches to an unwatermarked model or redesigns the sequence may remove it, and the absence of a watermark doesn't show that a sequence is natural or safe. Real use in DNA synthesis screening and databases will need synthesis companies, databases and model developers to coordinate on adopting a standard.
Scope note: experimental results come from DeepMind's blog and the Nature paper; the phage experiments are early results; for the licenses of the open-sourced parts, see the official repository.
via: Google DeepMind blog, Google blog, Nature paper, Nature news