Google DeepMind launches SynthID Bio, a family of methods that embeds detectable watermarks into AI-designed protein sequences and predicted three-dimensional structures. A peer-reviewed Nature paper reports that the signatures remained detectable while preserving protein function in laboratory tests.

 

The September 30 release extends provenance technology beyond digital media and into biological designs that can be physically synthesized. The work is a proof of concept, not a guarantee that a protein is safe or that every attempt to remove its watermark will fail.

 

The research establishes four main results:

  • Watermarks were added to protein sequences and predicted structures.
  • Watermarked binders remained functional against three targets.
  • Structure detection exceeded 99.8% at a 0.1% false-positive rate.
  • Code, laboratory data and model-weight access are being released.

 

More on This Story

 

Google DeepMind SynthID Bio Marks Sequences and Structures

Google DeepMind describes two related systems. SynthIDBio-sequence guides amino-acid choices while ProteinMPNN generates a protein sequence. SynthIDBio-structure fine-tunes part of AlphaFold 3 so its predicted atomic coordinates carry a statistical signature.

 

The sequence method adapts tournament sampling from SynthID's text-watermarking work. Candidates then pass through normal structural filters and an additional detectability threshold. The structure method instead trains the watermark into AlphaFold 3's diffusion and confidence modules, making the signal part of the model's generated coordinates.

 

Both are zero-bit watermarks. They can indicate that a design came from a participating process, but they do not encode a detailed message, identify an individual user or record every step in a design's history.

 

Nature Tests Protein Binders Against Three Targets

The Nature study evaluated protein binders aimed at VEGF-A, the SARS-CoV-2 spike receptor-binding domain and PD-L1. Researchers started from known AlphaProteo binder backbones, generated watermarked and unwatermarked sequences with ProteinMPNN, and measured binding in the laboratory.

 

The team obtained low-nanomolar watermarked binders for the viral target and subnanomolar binders for VEGF-A and PD-L1. Across the tested populations, watermarking did not produce a significant difference in binding-affinity distributions or in hit rates at the study's principal threshold.

 

Detection involved an explicit trade-off. Filtering produced a 100% true-positive rate for the synthesized sequence designs at a calibrated 0.1% false-positive rate, but stricter filtering meant generating more candidates. The paper says the extra practical cost was negligible because rejected candidates did not need the most expensive AlphaFold 3 validation stage.

 

For predicted structures, the fine-tuned model exceeded a 99.8% true-positive rate at the same false-positive level while maintaining structural-accuracy metrics. Training took roughly one day on 256 Nvidia A100 GPUs, and the watermarking method adds no inference overhead once the model is fine-tuned.

 

Protein Provenance Could Support Screening and Data Integrity

Provenance matters because AI tools can generate sequences that differ substantially from known natural proteins. Conventional similarity screening may therefore have less information to work with. An intact watermark could give DNA-synthesis providers another signal that a sequence was created through a recognized design system.

 

The researchers do not propose replacing existing screening. They describe a layered approach in which watermark detection complements customer checks, similarity analysis and other safeguards. A watermark identifies likely origin; it does not determine whether a sequence is harmless.

 

Scientific databases are a second use case. Protein Data Bank, UniProt and GenBank underpin model training and bioinformatics. Detectable provenance could help curators scrutinize synthetic entries before mislabeled data enter public repositories and distort later research.

 

Google DeepMind is open-sourcing the sequence code and in-vitro data through its SynthID Bio repository. The project also provides a route to request the recommended structure-model weights, subject to terms that prohibit several uses.

 

Resequencing and Adoption Remain Open Problems

The sequence watermark can be removed by running a design through ProteinMPNN again. In the paper's simulations, that attack reduced the estimated chance that the altered binder would retain useful function, but it did not make removal impossible. More robust defenses against partial or complete resequencing remain research problems.

 

The structure watermark survived noise, coordinate rounding, rigid transformations and some cropping. It was less robust to a process called structural relaxation, which adjusts atomic coordinates toward lower-energy configurations. The current detector also cannot distinguish marks from multiple providers or users.

 

Deployment will require synthesis companies, database maintainers and model developers to agree on detector access, secret-key handling and acceptable error rates. False negatives could miss synthetic material, while false positives could trigger unnecessary scrutiny or undermine trust in provenance claims.

 

SynthID Bio therefore represents a demonstrated building block rather than a completed biosecurity system. Its immediate achievement is narrower and measurable: an AI provenance signal can survive the jump from a generated sequence into a functioning laboratory-made protein without erasing the function researchers designed it to perform.