📥 Content Hub
← назад
AI / Искусственный интеллект dqindia.com en 2026-10-03 02:35 4 min

Google DeepMind develops watermark for AI generated proteins - dqindia.com

Кратко: Google DeepMind has introduced SynthID Bio, a family of techniques that embeds detectable signatures directly into AI-generated protein sequences and predicted three-dimensional structures. Unlike metadata attached to a digital file, the watermark is incorporated into the biological design itself.
🧭 Извлечение: ok · confidence 90% · диагностика
High confidence: full text extraction produced 5354 characters.

Google DeepMind has introduced SynthID Bio, a family of techniques that embeds detectable signatures directly into AI-generated protein sequences and predicted three-dimensional structures. Unlike metadata attached to a digital file, the watermark is incorporated into the biological design itself.

The underlying research, published in Nature on September 30, describes the technology as a proof of concept, an important distinction as Google explores its potential use in biosecurity and scientific integrity rather than presenting it as a complete safeguard. That is the problem SynthID Bio is trying to address.

Putting the watermark inside the protein

SynthID Bio uses different techniques for protein sequences and predicted structures.

For sequences, it subtly influences which amino acids are selected during protein generation. Google tested the approach using AlphaProteo designs alongside a SynthID Bio-enabled version of ProteinMPNN.

For predicted three-dimensional structures, the researchers fine-tuned part of AlphaFold 3's diffusion network so that its output contains a detectable signal within the atomic coordinates. Because the watermarking capability is incorporated into the model, the resulting structures carry the signal when generated by that modified model.

The critical question was whether embedding such a signature would alter the protein itself.

Google's wet-lab experiments tested protein binders against three targets: VEGF-A, the SARS-CoV-2 receptor-binding domain, and PD-L1. The company reported that watermarked designs showed comparable binding affinity and hit rates to their non-watermarked counterparts. The Nature paper similarly reports that watermarking did not affect the binding-affinity distribution or hit rates across the tested targets.

For structural predictions, the researchers reported near-perfect watermark detectability with negligible effects on overall prediction accuracy under the recommended configuration.

A biological watermark has little practical value if inserting it compromises the molecule's intended function.

AI is creating a new screening problem

DNA synthesis providers already screen customer orders against known biological threats. But generative models can create sequences with limited similarity to existing natural or hazardous sequences.

That makes provenance increasingly useful alongside conventional screening.

Instead of relying only on what an unfamiliar sequence resembles, a synthesis provider with access to the appropriate detection system could potentially determine whether a design carries a watermark associated with an AI model.

Google argues that this could provide another signal within a broader, layered biosecurity system. It could also help repositories such as the Protein Data Bank, UniProt, and GenBank distinguish appropriately identified synthetic entries from other biological data.

This makes SynthID Bio less a replacement for existing screening than an additional provenance layer.

But the watermark can be removed

This is also where the limitations become important. SynthID Bio is not an indelible molecular fingerprint.

The Nature paper finds that attempts to resequence proteins can largely remove the sequence watermark. The structural watermark can withstand some basic transformations and noise, but the researchers also found that structural relaxation can destroy it.

A separate Nature report on the research highlighted the same weakness: someone could potentially pass a watermarked protein through another design process and generate a sequence that obscures the original marker.

Google itself acknowledges that improving resistance to deliberate tampering remains a key challenge.

There are other constraints. The current system essentially identifies the presence of a watermark rather than encoding rich provenance information or distinguishing users. Operational deployment would also require coordination between AI developers, synthesis companies, biological repositories, and other trusted parties that can detect the signal.

So the technology should not be read as a way to declare an unfamiliar protein safe simply because a watermark is detected. Rather, it provides another piece of information about provenance.

From AI content to biological provenance

SynthID started as Google's approach to identifying AI-generated digital content. DeepMind says the technology is already used for generated images, audio, text, and video. SynthID Bio extends that underlying idea into synthetic biology.

The researchers are already testing the concept beyond individual proteins.

DeepMind says it has worked with the Hie lab at Stanford University and Arc Institute to integrate SynthID Bio with Evo 2 and watermark the genome of an AI-designed bacteriophage. Early laboratory testing found the watermarked bacteriophages remained functional, although Google says further technical details are still to come.

The company is also open-sourcing SynthID Bio's code and in vitro data, with access instructions for the structural model weights available to researchers.

The larger issue is likely to outlast any individual watermarking technique. As AI moves from analysing existing biology to creating biological sequences that may never have existed in nature, provenance becomes harder to establish using traditional databases alone.

Читать оригинал ↗

Сделать контент из этого материала