Skip to content
MCAP $2.93T ▼-2.34%
BTC $85,749 ▲+0.72%
ETH $2,704 ▲+0.64%
BNB $783.56 ▼-0.28%
XRP $1.500 ▲+1.02%
SOL $120.81 ▲+1.46%
DOGE $0.0953 ▲+1.41%
ADA $0.273 ▲+4.08%
TRX $0.3362 ▼-0.05%
LINK $13.95 ▲+1.39%
AVAX $11.45 ▲+5.54%
HYPE $92.29 ▼-0.15%
DOT $1.210 ▲+1.27%
AI & crypto · Intermediate

Data provenance for AI explained: how blockchains prove where data and model outputs came from

How C2PA credentials, hashes, on-chain attestations, XYO's proof of origin, ERC-8004 and zkML prove where AI data and outputs came from, what the EU AI Act requires from August 2026, and what a verifier actually checks.

Crypto Coin Show Editorial Desk·Updated September 28, 2026·24 min read·Educational, not investment advice

Key takeaways

  • From August 2, 2026 the EU AI Act’s Article 50 requires generative AI providers to mark outputs in a machine-readable way and deployers to disclose deepfakes, with fines of up to EUR 15 million or 3 percent of global turnover, per the Commission’s June 2026 Code of Practice and legal commentary.
  • C2PA Content Credentials became the default format for signed provenance metadata: OpenAI joined the C2PA steering committee and added Google’s SynthID watermark to its image outputs on May 19, 2026, and the coalition counted more than 5,000 participants as of June 2026.
  • Blockchains add what a signed file cannot: a neutral, timestamped, append-only record that a hash existed at a point in time. Ethereum Attestation Service alone reported more than 9.5 million attestations from over 450,000 attesters as of October 2026.
  • Provenance is moving from media files to AI agents. ERC-8004, the Trustless Agents standard co-authored by MetaMask and Ethereum Foundation staff, went live on Ethereum mainnet in late January 2026, and XYO wrote its first independent proof-of-performance records for Theta EdgeCloud agents in May 2026.
  • Verifiable inference (zkML) is real but early: EZKL can prove ONNX models with a modified Halo2 prover, yet no major AI lab proves production outputs today, while Deloitte projected US generative AI fraud losses of USD 40 billion by 2027, up from USD 12.3 billion in 2023.

Who this is for: Newsroom and AI lab operators, compliance leads, infrastructure investors and policy staff who need to decide what “provenance” should mean in their own pipeline, and who want to tell a signed file from an on-chain attestation from a zero-knowledge proof before a vendor does it for them.

Every piece of digital content now arrives with the same question: where did this come from, and has it been changed? A photo from a protest, a dataset used to fine-tune a model, the answer an AI agent gives before it moves money. None of them carries evidence of its own history. Pixels do not remember the camera that captured them, and a model’s output does not remember the weights or prompt that produced it.

Three forces made this urgent in 2025 and 2026. Deepfakes became a fraud tool: Deloitte projected in May 2024 that generative AI could push US fraud losses to USD 40 billion by 2027, from USD 12.3 billion in 2023. Training data disputes moved into courtrooms, and labs began paying for licensed, consented data. And regulators acted: the EU AI Act’s Article 50 transparency obligations started applying on August 2, 2026, with a Code of Practice finalised on June 10, 2026 and roughly 190 signatories by late July.

This guide explains the toolkit that emerged in response: C2PA signed metadata, hashing and signatures, attestation networks such as Ethereum Attestation Service, data-focused chains like XYO Layer One, IP and data marketplaces, decentralised storage, zero-knowledge proofs of inference, and proving what an AI agent did. We are specific about what a blockchain adds over a signed log, where each approach breaks, what it costs and what a verifier actually checks. It pairs with our guide on DePIN, since much real-world provenance depends on physical devices.

AI data provenance by the numbers

Aug 2, 2026EU AI Act Article 50 transparency rules applyEuropean Commission, June 2026
EUR 15m / 3%Maximum fine for Article 50 breachesMishcon de Reya, June 2026
5,000+C2PA participants and affiliatesContent Authenticity Initiative, June 2026
9.5m+Attestations recorded via EASattest.org, October 2026
10m+Devices in XYO’s data-collecting networkMarkus Levin, CCS interview, Aug 2026
$40bnProjected US gen-AI fraud losses by 2027Deloitte, May 2024

The problem: four provenance gaps

“Provenance” is one word covering four questions, and the tools differ for each.

1. Is this media real?

Synthetic images, voices and video are cheap enough that detection by eye is unreliable and detection by classifier is an arms race the detectors tend to lose. The response has shifted from detecting fakes to proving authenticity: the verifier asks “can anyone prove this is real?” and treats unsigned content as unverified rather than false.

2. What was this model trained on?

Labs face copyright litigation, consent disputes and, under the EU AI Act’s general-purpose model rules in force since August 2025, a duty to publish a summary of training content. This gap pushed Story Protocol to rebrand as The Data Foundation in June 2026, telling users that “the form of IP pulling hardest was AI training data.”

3. What did this agent do?

AI agents now call APIs, move funds and book resources for other software. When one misbehaves, the parties need a record of what it saw, decided and when. Internal logs are controlled by the agent’s operator, exactly the party a counterparty may not trust. ERC-8004 on Ethereum and XYO’s “proving action” both target this gap.

4. Can a regulator check any of this later?

Article 50 turns the first gap into a legal duty. Providers must mark AI outputs in a machine-readable format; deployers must disclose deepfakes and AI-generated text on matters of public interest. Legacy generative systems on the EU market before August 2, 2026 have until December 2, 2026 for the marking duty, but the deepfake disclosure duty applied from August 2 with no grace period, per Baker McKenzie’s July 2026 analysis.

Approach one: signed metadata and C2PA

The Coalition for Content Provenance and Authenticity (C2PA) was co-founded in February 2021 by Adobe, Arm, BBC, Intel, Microsoft and Truepic, building on Adobe’s Content Authenticity Initiative from November 2019. Its output, Content Credentials, is a manifest embedded in a media file holding assertions (who created the asset, with what tool, what edits), a hash of the content and a digital signature from the device or software making the claim. Each later editor appends its own signed assertion, so the manifest becomes a chain: camera, editing software, publisher’s CDN.

Adoption accelerated through 2025 and 2026. Cloudflare announced on February 3, 2025 that Cloudflare Images preserves Content Credentials through resizing and signs its own transformations as new manifest actions. Google’s Pixel 10 shipped with C2PA photo credentials, and by May 2026 Google extended C2PA to video on Pixel 8, 9 and 10. The largest step came on May 19, 2026, when OpenAI joined the C2PA steering committee alongside Adobe, Microsoft and Google and began embedding Google DeepMind’s SynthID watermark in images from ChatGPT, its API and Codex, with Kakao, ElevenLabs and Nvidia announcing adoption the same week. The CAI counted more than 5,000 participants as of June 2026.

Why pair a watermark with metadata? Because they fail differently. C2PA metadata is rich but fragile: a screenshot, a re-encode or a platform that strips metadata on upload removes it. SynthID is an invisible statistical watermark in the pixels that survives screenshots and compression but says little beyond “a Google or partner model made this.” The EU’s final Code of Practice reflects the same logic, asking providers for a layered approach combining machine-readable marking, metadata, imperceptible watermarking and fingerprinting.

What C2PA does not do

A valid Content Credential proves that a specific signer made a specific claim about a file at a specific time. It does not prove the claim is true: a camera certificate attests that the firmware signed the image, not that the scene was real. The trust model is a certificate hierarchy like the web’s TLS, so a verifier must decide which signers to trust. And a missing credential is not evidence of forgery; most content online is unsigned and will remain so for years, which is why the standard works best as a positive signal rather than a filter. Critics also raise privacy concerns, since manifests can carry a lot of metadata about who signed what.

Approach two: hashing, signing and what a blockchain adds

Underneath every provenance system sit two primitives. A cryptographic hash (SHA-256 is the norm) turns any file into a fixed-length fingerprint; change one byte and the fingerprint changes completely. A digital signature binds a statement to that fingerprint so anyone with the public key can check it came from that key unaltered. C2PA packages these inside the file. A plain signed log is another packaging: a company writes “hash H of dataset D existed at time T” into a database and signs each line.

So why is a blockchain needed? Because a signed log proves integrity, not timing or completeness, and it is controlled by the party making the claim. The owner can delete an entry, insert a back-dated one, or show different versions to different auditors. A signature says who signed, not when, and not whether other entries once existed. A public blockchain fixes exactly these weaknesses and nothing else.

Property Signed file (C2PA) Private signed log Public chain attestation Zero-knowledge proof
Proves content unchanged Yes, via hash Yes, via hash Yes, via hash Yes
Proves who made the claim Yes, certificate chain Yes, key Yes, wallet or key Yes, via public inputs
Proves when the claim existed Signer’s own clock Operator’s own clock Block timestamp, independent Only if anchored on-chain
Resists deletion or back-dating No No Yes Yes, if anchored
Proves the claim is true No No No Yes, for the computation proven
Survives screenshots and re-encoding No (watermark needed) n/a Yes, hash stored separately Yes
Typical cost per record Near zero Near zero Cents on L2s, more on Ethereum L1 Seconds to hours of compute

A public chain is a neutral notary with a clock. It does not add truth: anchor the hash of a fabricated dataset and the chain faithfully proves the fabrication existed at that time. Serious designs therefore pair a chain with evidence produced at the source: a device signature, independent witnesses, or a proof of computation.

Approach three: on-chain attestation networks

Ethereum Attestation Service (EAS) is the most widely used general-purpose attestation layer. An attestation is a signed statement by an attester about a subject, structured by a schema anyone can register, published on-chain or kept off-chain and revealed on demand, and revocable. As of October 2026 EAS reported more than 9.5 million attestations from over 450,000 attesters across Ethereum, Base, Optimism, Arbitrum, Polygon and other networks, with Coinbase using it for identity verifications, and it now lists AI model evaluations and data oracles among its use cases.

For AI provenance, a schema might read: dataset hash, licence identifier, source organisation, consent reference, capture date. A lab attests to that tuple on ingest, a licensor attests that it granted the licence, and an auditor later attests that it checked both. None needs to trust the others’ databases, and all three records are timestamped by the chain.

Data-native chains: XYO Layer One

XYO Network works on a different piece of the puzzle: making a claim about the physical world trustworthy at the moment it is created. Its core primitive is the bound witness, in which two or more interacting devices co-sign a record of that interaction, so each becomes evidence for the other. Proof of Origin chains those records together, and XYO’s witness-based consensus discounts outliers: as co-founder Arie Trouw put it in an April 2026 Crypto Coin Show interview, if four of five sources agree and the fifth does not, the fifth is probably the problem. XYO’s device network counted more than 10 million nodes as of August 2026 according to co-founder Markus Levin.

XYO Layer One, launched in Q3 2025, is a blockchain built for data rather than payments. It stores metadata, hashes and timestamps on-chain and keeps datasets in public or private Data Lakes, which Trouw calls reveal-on-demand privacy. XYO is staked to operate the network and XL1 pays for transactions, with a portion of each payment burned. April 2026 upgrades raised throughput two to five times; in May 2026 XYO launched an AI SDK and Data Lakes with about 2,000 early-access signups; Crypto.com listed XL1 and XYO in July 2026; and on September 1, 2026 Gate AI and XYO said they were exploring AI-powered crypto cards verified on the chain. Our conversation with Levin on XYO’s accountability layer for AI agents covers the design in depth.

Approach four: IP, data markets and storage

If attestations answer “who said what about this data,” a second group of projects makes the data itself a licensable, payable asset.

Story Protocol becomes The Data Foundation

Story Protocol launched as a chain for registering IP as programmable assets with licence terms attached. On June 25, 2026 it announced it was becoming The Data Foundation, converting the IP token one-for-one to DATA, because AI training data had become the dominant form of IP registered. Its stack now includes Trace, a public ledger where labs audit a dataset’s consent and payment records by dataset, modality and time; Kled, an opt-in human-data marketplace reporting more than 1.1 billion uploaded files at about 5 million a day as of June 2026; and Poseidon, which filters out scraped, synthetic or altered content. As of October 2026 its homepage showed 267 million receipts and 529,000 contributors. Provenance stopped being a feature of IP registries and became the product.

Ocean, Vana and Numbers

Ocean Protocol tokenises datasets as data NFTs with ERC-20 datatokens granting access, and its Compute-to-Data runs models against data that never leaves the owner’s environment, so “this model trained on this dataset” is enforced at the point of compute. Vana organises user data into DataDAOs and pays contributors when models train on it; its Vega upgrade went live on September 22, 2026 with data reads of about 50 milliseconds, followed by expanded staking and a public dashboard on September 28, 2026. Numbers Protocol bridges media and data: its Jade mainnet, launched December 2022, registers each asset under a Numbers ID with on-chain capture and edit history and supports both IPTC and C2PA metadata, so a C2PA-signed photo also gets an independent on-chain timestamp. Its documentation cites Reuters’ 2020 US election coverage as an early use case.

Where the bytes live

A hash on a chain proves a file existed; it does not keep the file available. Walrus, built by Mysten Labs on Sui, erasure-codes files across nodes, charges in WAL and in 2026 markets “Walrus Memory” for AI agents. Autonomys, on mainnet since Q4 2024, sells one-time-payment permanent storage through Auto Drive; its October 2026 snapshot showed 182 nodes, 21.71 petabytes pledged, 256x replication and 0.27 AI3 per megabyte. Filecoin remains one of the largest incentivised storage networks and the usual archive behind IPFS content identifiers, which are themselves hashes and therefore natural provenance anchors. The practical pattern is hybrid: originals in cloud storage, a hash and attestation on a public chain, and an optional decentralised replica for independence from any single vendor.

Verifiable inference: zkML and its real maturity

Everything above proves things about inputs and records. Zero-knowledge machine learning (zkML) tries to prove the computation itself: that a specific model, identified by a commitment to its weights, produced a specific output from a specific input, without revealing the weights or input. At scale, a verifier could check that a lab’s published model, not a cheaper substitute, generated an answer, or that an agent’s decision came from the policy it claimed.

EZKL, from zkonduit, is the most accessible toolchain. It compiles ONNX models into circuits and generates proofs with what its documentation calls a highly improved version of Halo2, the proving system developed by Zcash, with a hosted cluster called Lilith for heavier jobs. Modulus Labs built specialised provers it describes as 1,000 times more capable than legacy approaches, framing the goal as “Accountable AI” for smart contracts. Giza began in zkML and by 2026 had pivoted to verifiable DeFi agents; its ARMA product on Base was listed as coming soon as of October 2026, with its dashboard showing zero assets under agent management.

The real state of the field: proving small and medium models, decision trees and compact networks, is practical today and used in a handful of on-chain applications. Proving a frontier language model’s inference per output is not; proofs for large transformers still take orders of magnitude longer than the inference itself, with costs in GPU time rather than cents, and no major AI lab proves production outputs with zkML as of October 2026. Nearer-term uses are narrower: proving a published checkpoint matches a committed hash, or that an agent followed a stated policy for one high-value action. Space and Time applies the same idea to data, using Proof of SQL to prove a query result was computed correctly against a committed dataset; it is live on mainnet with SXT staked by operators and lists Microsoft, Nvidia and Chainlink among partners as of October 2026.

Agent action provenance: ERC-8004 and proving action

ERC-8004, titled Trustless Agents, was co-authored by Marco De Rossi, head of AI at MetaMask, and Davide Crapis, AI lead at the Ethereum Foundation, among others, and went live on Ethereum mainnet around January 30, 2026. It defines three registries: an Identity Registry giving each agent a portable, NFT-compatible identifier; a Reputation Registry for signed client feedback; and a Validation Registry through which an agent asks a validator contract to check its work and record the result on-chain. Trust is tiered by risk, so low-stakes tasks rely on reputation while high-stakes ones require validation. QuickNode’s explorer showed agent identifiers above 23,000 on mainnet in the first week of February 2026, though many are unconfigured test registrations and no live total is published.

XYO’s approach, which it calls proving action, puts an independent observer outside the system being measured. On May 28, 2026 XYO and Theta announced an integration in which XYO nodes observe Theta EdgeCloud AI agents from outside, measure uptime, speed and reliability, and write a proof-of-performance record to XYO Layer One, with raw observations in Data Lakes anyone can query; each record is a transaction paid in XL1. EdgeCloud agents were already running workloads for the Houston Rockets and Olympique de Marseille. The design principle, in XYO’s words, is that the data never touches the system it is measuring, which addresses the core weakness of self-reported logs. Our June 2026 interview with Markus Levin covers the rationale.

The two are complementary: ERC-8004 is a registry layer any validator can plug into, and XYO is one validator design with its own chain and witness network. Of any such system, ask who observes the agent, whether the observer has an incentive to lie, and whether the record can be deleted after the fact.

Regulation: the EU, US states and the NO FAKES Act

The EU AI Act entered into force on August 1, 2024 with a staggered timetable: prohibited practices from February 2025, general-purpose model duties from August 2025, most remaining obligations including Article 50 from August 2, 2026, and some high-risk rules from August 2027. The Commission published draft Article 50 guidelines on May 8, 2026, the final Code of Practice on June 10, 2026, and a set of labelling icons. About 190 organisations had signed by late July 2026 while the Commission assessed the Code’s adequacy. Signing is voluntary, but adherence is the intended route to demonstrating compliance. The Code names no standard, but its language on digitally signed metadata, imperceptible watermarking and fingerprinting maps closely onto the C2PA plus SynthID pattern the industry had already adopted.

The United States has no federal provenance mandate. The National Conference of State Legislatures recorded that all 50 states introduced AI legislation in 2025 and 38 states enacted around 100 measures, many targeting intimate and election deepfakes. California’s AB 853, pending during 2025, would require large platforms to retain provenance data and device makers to build provenance into capture devices, effectively mandating C2PA-style credentials in hardware; check current status, since several bills were still moving at our last verification.

The NO FAKES Act, which would create a federal right over a person’s voice and likeness against unauthorised digital replicas, was introduced in July 2024 and reintroduced in April 2025 as S. 1367 with endorsements from SAG-AFTRA, Universal Music Group, OpenAI, Google, Disney and Adobe; the Electronic Frontier Foundation and the Center for Democracy and Technology oppose it on speech grounds. As of the most recent record we could verify, December 2025, it remained in the Senate Judiciary Committee without a floor vote in either chamber, and we could not confirm any 2026 action. For how US agencies divide authority over related digital assets, see our guide to SEC versus CFTC crypto regulation.

Enterprise pilots and what they reveal

Adopters cluster into three groups: media and platform companies on signed media (Cloudflare since February 2025, Google across Pixel capture, OpenAI, Nvidia, ElevenLabs and Kakao since May 2026); sports and infrastructure businesses on agent provenance through Theta EdgeCloud agents for the Houston Rockets and Olympique de Marseille; and AI data buyers on audited pipelines such as The Data Foundation’s Trace ledger. The shared lesson: provenance only works when attached at creation and preserved through every hop, so the first enterprise job is often stopping systems from stripping it.

How we got here: a timeline

C2PA formed. Adobe, Arm, BBC, Intel, Microsoft and Truepic co-found the coalition behind Content Credentials.

EU AI Act enters into force. Article 50 transparency duties are scheduled for August 2026.

Cloudflare preserves Content Credentials in Cloudflare Images. Provenance survives the CDN at scale for the first time.

XYO Layer One launches. A chain built for data rather than payments, with XL1 as gas.

EU publishes first draft Code of Practice on AI-generated content. A second draft follows on March 3, 2026.

ERC-8004 Trustless Agents goes live on Ethereum mainnet. Identity, reputation and validation registries for AI agents.

OpenAI joins the C2PA steering committee and adopts SynthID. XYO and Theta announce proof-of-performance for AI agents on May 28.

EU finalises the Code of Practice; Story becomes The Data Foundation. Published June 10 and announced June 25 respectively.

Article 50 obligations apply. Legacy generative systems get until December 2, 2026 for machine-readable marking only.

Vana ships the Vega upgrade. Data reads in about 50 milliseconds, with expanded staking on September 28.

Worked example: a newsroom and a lab attach provenance end to end

Consider a news organisation publishing 300 photographs and 40 video clips a day that also licenses a 50,000-image archive to an AI lab for fine-tuning. Both want provenance a third party can check. Costs are illustrative, based on public 2026 pricing patterns, not quotes.

  1. Capture. Photographers shoot on C2PA-capable devices, which sign a manifest with capture time, device certificate and a SHA-256 hash of the image. Marginal cost: zero.
  2. Edit. Editing software appends a second signed assertion listing the actions. The original hash stays inside the manifest so the edit lineage is visible.
  3. Anchor. At publication a script writes an attestation of (asset hash, manifest hash, URL, timestamp) to an L2 such as Base via EAS. At typical 2026 L2 gas costs of a fraction of a cent to a few cents per transaction, 340 attestations a day is under USD 10 a day at the high end, or under USD 4,000 a year. Batching hashes into a daily Merkle root cuts that to one transaction a day.
  4. Deliver. Images go through a CDN with Content Credential preservation enabled, and a public verification page re-hashes any downloaded image and checks it against both the manifest and the attestation.
  5. License. The newsroom computes a Merkle root over all 50,000 image hashes and attests (archive root, licence terms hash, licensee, date) on-chain. The lab recomputes the root from the files it received and publishes its own attestation referencing the newsroom’s. Two parties, two signatures, one shared root.
  6. Train and disclose. The lab records the archive root in its model card and EU training-content summary. Generated images carry a C2PA manifest naming the model plus a watermark, meeting the Article 50(2) marking duty from August 2, 2026.
  7. Dispute. A year later a photographer claims an image was used without licence. The verifier hashes the disputed image, checks whether that hash is a leaf under the licensed root, confirms the attestation’s block timestamp predates the training date, and checks that both parties’ keys signed. If the hash is not in the tree, the lab cannot have received it under the licence; if it is, the licence terms hash settles what was permitted.

Total incremental spend for the newsroom: a few thousand dollars a year in gas plus engineering time. What the verifier never has to do is trust either party’s internal database, which is the point.

How to evaluate a provenance system: a checklist

  • Where is the evidence created? At capture (device signature, bound witness) is strongest; at upload is weaker; “we hash it later” is a log, not provenance.
  • Who controls the record after it is written? If the claimant can delete or reorder entries, the system proves integrity but not history. Look for public chain anchoring or multi-party attestation.
  • Does metadata survive the delivery path? Push a real image through your CDN, CMS and social platforms. Where platforms strip metadata, you need preservation or a separate on-chain hash.
  • Is there a watermark for when metadata is lost? The EU Code asks for layered marking. Metadata alone fails on screenshots.
  • Can a verifier check it without the vendor? Open tools (contentcredentials.org/verify, a public explorer, open-source proof verifiers) beat a vendor dashboard.
  • Is “verified” being confused with “true”? A valid signature proves the signer’s claim, not reality. Ask how the system handles a trusted signer that lies.
  • For agents: who observes the agent? Self-reported logs are weakest; independent observers (XYO’s model) or on-chain validators (ERC-8004) are stronger.
  • What does it cost at your volume? Model cents per record on an L2, compute hours for proofs and storage fees in a volatile token. Budget in fiat.

Risks and open questions

The oracle problem never goes away. Every system here reduces to trust in whoever produced the first signed claim. A compromised camera key, a bribed witness node or a mislabelled dataset propagates through every attestation built on it. Redundancy and independent observers reduce the risk but cannot remove it, and liability when a trusted issuer is wrong remains unsettled.

Fragmentation and privacy. C2PA won the metadata layer, but social platforms still strip or ignore credentials inconsistently, and Google’s promised browser verification was not live as of May 2026. On-chain records are scattered across EAS on a dozen networks, XYO Layer One, Numbers’ Jade and others, with no common way to resolve a hash across them. Immutable attestations about people also collide with GDPR erasure rights; keeping personal data off-chain and anchoring only hashes mitigates this, but a hash of a known file is still a permanent link to it.

Token economics versus utility. Autonomys’ 0.27 AI3 per megabyte, XYO’s XL1 burn per record and Walrus’ WAL fees are all exposed to token prices that swing far more than the service’s value. Enterprises want fiat invoices, and intermediaries absorbing token risk will take a margin. It is also an open question whether attestation needs its own chain, or whether EAS-style contracts on existing Ethereum Layer 2s capture most of the value.

What to watch next

  • December 2, 2026: end of the Article 50(2) grace period. Legacy generative systems in the EU must mark outputs machine-readably from this date; expect a wave of C2PA and watermark rollouts in Q4 2026.
  • Commission adequacy decision on the Code of Practice (late 2026). Endorsement would make the Code the de facto compliance route and shape which technical methods count.
  • Google Search and Chrome C2PA verification rollout. Promised in May 2026; browser-level display of credentials would change user expectations quickly.
  • NO FAKES Act movement in the 119th Congress. A Senate Judiciary markup in late 2026 or early 2027 would be the first formal action since April 2025.
  • California AB 853 and device-level provenance mandates. If enacted, capture-device requirements would push C2PA into hardware by law.
  • Agent registries moving from registrations to validations. Watch whether ERC-8004 validation activity and XYO proof-of-performance volumes in 2027 reflect production use rather than test entries.

Glossary

Provenance
The documented history of data or content: who created it, when, with what, and what changed afterwards.
Hash
A fixed-length fingerprint computed from a file, such that any change to the file changes the fingerprint. SHA-256 is the common choice.
Digital signature
A value created with a private key that lets anyone with the public key confirm a statement came from the key holder unaltered.
C2PA / Content Credentials
An open specification for embedding a signed manifest of capture and edit history inside a media file.
Watermark
A signal hidden in the content itself that survives copying and compression. SynthID is Google’s watermark.
Attestation
A signed, structured statement by one party about a subject, often recorded on a blockchain so its timing can be checked independently.
Bound witness
XYO’s primitive in which two or more interacting devices co-sign a record of the interaction, so each claim is backed by the other.
Proof of Origin
XYO’s chaining of bound witness records so a device’s observation history can be traced and verified.
Merkle root
A single hash computed from a tree of many hashes, allowing proof that one item belongs to a set without publishing the whole set.
zkML
Zero-knowledge machine learning: proofs that a specific model produced a specific output, verifiable without rerunning the model.
Deepfake
Under the EU AI Act, AI-generated or manipulated content resembling real people, places or events that would falsely appear authentic.

Why it matters

For a decade, the crypto industry argued that public ledgers could serve as neutral infrastructure for facts no single party should control. Payments were the first test. Data provenance for AI may be the second, and it arrives with a regulatory tailwind payments never had: from August 2026 the world’s largest trading bloc legally requires machine-readable evidence of what AI produced, and the fines are material. The pieces that work today are modest and specific, a signed manifest, a hash on a chain, an independent witness. The transformative piece, proofs of inference for large models, is still years from routine use.

For operators, the takeaway is to attach provenance at the source, protect it through every hop, treat chains as a notary rather than a truth machine, and budget for verification a stranger can perform without your help. For investors, the signal is not token prices but whether verification volume on these networks reflects real content and real agents, and whether enterprises are paying for it in fiat.

Sources

  1. European Commission: Code of Practice on Transparency of AI-Generated Content, June 10, 2026
  2. Baker McKenzie: New EU guidance on AI transparency from 2 August 2026, July 2, 2026
  3. Mishcon de Reya: AI Act transparency obligations, Code of Practice and draft guidelines, June 25, 2026
  4. National Conference of State Legislatures: Artificial Intelligence 2025 Legislation, 2025
  5. Wikipedia: NO FAKES Act, accessed October 6, 2026
  6. TNW: OpenAI adopts C2PA standard and Google’s SynthID, May 19, 2026
  7. Wikipedia: Content Authenticity Initiative, accessed October 6, 2026
  8. Cloudflare: Preserve Content Credentials with Cloudflare Images, February 3, 2025
  9. Deloitte: Deepfake banking fraud risk on the rise, May 29, 2024
  10. Ethereum Attestation Service: attest.org, accessed October 6, 2026
  11. The Block: Ethereum to launch standard for AI agent economy on mainnet this week, January 28, 2026
  12. QuickNode: ERC-8004 Explorer, Ethereum mainnet agent 23070, accessed October 6, 2026
  13. Crypto Coin Show: XYO Layer One achieved 2-5x throughput increase, Arie Trouw explains, April 20, 2026
  14. Crypto Coin Show: XYO and Theta Just Solved AI’s Accountability Problem, May 28, 2026
  15. Crypto Coin Show: How XYO is building an audit trail for AI agents, Markus Levin, August 14, 2026
  16. The Data Foundation: We’re becoming The Data Foundation, June 25, 2026
  17. Vana: Vega upgrade updates, September 28, 2026
  18. Numbers Protocol: Documentation introduction, accessed October 6, 2026
  19. Autonomys Network: network snapshot and Auto Drive, accessed October 6, 2026
  20. EZKL: Documentation, accessed October 6, 2026
  21. Modulus Labs: Accountable AI, accessed October 6, 2026
  22. Space and Time: Proof of SQL and mainnet, accessed October 6, 2026

Disclosure: XYO Network is a sponsor of Crypto Coin Show. They had no input into this guide. This guide is for education only and is not investment, legal or tax advice.

Frequently asked questions

What does data provenance mean for AI?

Data provenance is the verifiable history of a piece of data or an AI output: who created it, when, with what tool or model, and what changed afterwards. For AI it covers three things at once: whether media is authentic, what a model was trained on and under what licence, and what an autonomous agent actually did. The tools range from signed metadata such as C2PA to on-chain attestations and zero-knowledge proofs of computation.

What does a blockchain add that a signed log does not?

A signed log proves a record was not altered after signing, but the party that keeps the log controls it and can delete, reorder or back-date entries and show different versions to different auditors. A public blockchain supplies an independent timestamp, an append-only record and public verifiability. It does not prove a claim is true; it proves the claim existed in that form at that time.

What is C2PA and who uses it?

C2PA is the Coalition for Content Provenance and Authenticity, co-founded in 2021 by Adobe, Arm, BBC, Intel, Microsoft and Truepic. Its Content Credentials embed a signed manifest of capture and edit history in a media file. Google Pixel phones, Cloudflare Images and, since May 19, 2026, OpenAI's image products carry or preserve them, and the coalition counted more than 5,000 participants as of June 2026.

What does the EU AI Act require from August 2026?

Article 50 applied from August 2, 2026. Providers of generative AI must mark outputs in a machine-readable format and make them detectable, and deployers must disclose deepfakes and AI-generated text on matters of public interest. Legacy systems already on the EU market have until December 2, 2026 for the marking duty only. Breaches can draw fines of up to EUR 15 million or 3 percent of global annual turnover.

How does XYO prove where data came from?

XYO uses bound witnesses, in which two or more interacting devices co-sign a shared record of that interaction, and chains those records together as Proof of Origin. Its witness-based consensus discounts outlier sources. XYO Layer One, live since Q3 2025, stores hashes, metadata and timestamps on-chain while full datasets sit in public or private Data Lakes, and since May 2026 it has written proof-of-performance records for AI agents on Theta EdgeCloud.

Is zkML ready to verify AI model outputs?

Partly. Toolchains such as EZKL can generate zero-knowledge proofs that a model expressed in ONNX produced a given output, and this is practical for small and medium models today. Proving each inference of a frontier language model remains far slower and more expensive than the inference itself, and no major AI lab proves its production outputs this way as of October 2026. Nearer-term uses are checkpoint commitments and single high-value agent decisions.

What is ERC-8004?

ERC-8004, called Trustless Agents, is an Ethereum standard that went live on mainnet around January 30, 2026. It defines an Identity Registry giving AI agents portable identifiers, a Reputation Registry for signed client feedback, and a Validation Registry through which an agent can have its work checked by a validator contract and the result recorded on-chain. It was co-authored by staff from MetaMask and the Ethereum Foundation.

What does a verifier actually check in a provenance dispute?

A verifier recomputes the hash of the disputed file, checks that hash against the signed manifest and against any on-chain attestation or Merkle root, confirms the block timestamp predates the event in question, and verifies that the expected keys signed. If the hash is absent from the committed set, the content was not part of the licensed or published record; if present, the attached licence terms settle what was permitted.

This explainer is reviewed and updated as the rules and the market change. Last reviewed September 28, 2026. It is educational content and not financial, legal or tax advice.

Keep learning