Markus Levin / XYO Network
Markus Levin on XYO's data provenance layer for AI agents
In this episode
AI hallucinations aren't a model problem — they're a data problem. When AI systems have no way to verify what's real in the physical world, they fill the gaps with guesses. And when AI agents act on those guesses autonomously, there's no record of what happened or why. DATA PROVENANCE is the missing layer, and XYO just shipped it.
Markus Levin, Co-Founder of XYO, breaks down two major announcements: Data Lakes — now live at xyo.network/data-lakes — and a partnership with Theta Network that pairs XYO's verifiable on-chain data infrastructure with Theta's decentralized media and delivery layer. Markus explains why AI hallucinations are fundamentally a data provenance problem, what the absence of an audit trail means as AI agents become financial managers, logistics coordinators, and healthcare decision-makers, and how XYO's AI SDK lets developers add cryptographic proof to agent decisions and model outputs today.
- AI hallucinations stem from data quality issues rather than model problems, requiring verifiable data sources to validate AI outputs.
- XYO's Layer 1 blockchain is purpose-built for data storage using data lakes and cryptographic proofs, unlike general-purpose blockchains designed for transactions.
- Data provenance through proof of origin enables tracking where data originated and ensures immutability for autonomous AI agents and smart city systems.
- Storing large data quantities directly on blockchain remains expensive across Bitcoin, Ethereum, and Solana, making off-chain data lakes with hashed verification more practical.
- Decentralized physical infrastructure networks with verified data are essential as AI agents assume roles in healthcare, logistics, and autonomous vehicle coordination.
Chapters
Transcript
Read the full transcript
I'm Ashton Addison from the Crypto Coin Show, and today on Blockchain Interviews, we're back with us Marcus Levin, co-founder of XYO. XYO just made two major announcements with their data lakes and partnership with Theta, one of the longest-standing DePIN networks with over 10 million nodes and fully secured by the XL1 XYO Layer 1 network. Excited to dive into DePIN, AI, commerce, space,
AI hallucinations, AI agents and commerce and where everything is going in the merge between AI and decentralized networks. And there's so much to cover in space. We might even cover some of these other hot topics right now. Appreciate you taking the time, Marcus.
Yeah, thanks for having me again, Ashton. It's it's great to have your support. Thank you.
Yeah, you're very welcome. I've been a long-stand supporter of XYO for over 8 years since the launch, and you and the team were early to decentralized physical infrastructure networks. Now, the world is catching up and AI is making it catch up as physical AI comes into play. AI agents are moving into everyone's computers and houses, and we need to make sure that the data is
verified, secure, and people aren't looking into attacking your agents or going into your the robot in your house and seeing what you're doing in the in the privacy of your own space. And I know that decentralized networks and provenance of data can can solve that. I would love for you to touch on that a little bit and sort of the goal of XYO as a DePIN network in
in securing data and why it's important.
Yeah, you know, like our world runs on data, you know, anything you process, you like our brain runs on data, right? And then from there, if you the economy runs on data, and the world of the future, if you speak about AI and robots, that tons of data, right? It's It's the most important input. It's
the fuel which makes it run. If that data is corrupted or wrong, then then you know, that world doesn't doesn't work as well, you know? AI hallucinates, robots might do the wrong thing, you know, might even kill the wrong person, you know, in a war. And it can be impactful, or it does the wrong thing in a in a surgery center, right? As a as a doctor.
And so, we have to make sure that we have good data feeding that, because our world will be more and more automated. Now, if you think about smart cities, for example, right? They will run by themselves, you know? I see a world where we don't have traffic lights anymore, and cars are just self-driving, you know? And the city optimizes it for its
traffic flow, and you know, where the cops are, and everything, really. And so, for that to run well, you need data you can trust, and which is decentralized, and which is not easily hackable, and otherwise, you have problems. And we at XYO are able to provide that.
Yeah, I'm I'm looking forward to not having traffic lights, and everything
just moves perfectly smoothly with automated cars, which we're already at the beginning stages of it. We have it in a few cities, a few cars, but we're getting there. But, it seems more prominently right now, everyone is pretty familiar with AI chatbots, some working with AI agents as well. There's definitely some issues with the truth of the information, maybe
where the source and where it comes from, hallucinations. If you can explain that and how that might impact everyday people that are in the office or at home using AI and why decentralized networks really need to be a part of this.
Yeah. Yeah, AI is so amazing. Like it has changed our lives in the last 4 years, you know, like for some slower than expected, for
some faster than expected. I think Elon thought we are at AGI right now, you know, general intelligence of AI. But it's it's it's we're not quite there yet, right? But it definitely had a big impact. And we are now, and, and for that, you know, to run well, we need that good data. And we need to be able, like, for example,
for an AI, if an AI hallucinates, we need to understand, why. And that, you know, if an AI gets a feedback, it says, "Okay, I got these three data sources, which I based my answer on. And then, we would be able to validate, those data sources. And then we must say, "Okay, this one is wrong, for example."
Mhm.
So, what we're doing with XY O is we
have a huge network of more than 10 million devices which collect data. It collects and validates it, through different kinds of algorithms and cryptographic proofs. And then that data gets put onto the XY O layer one and its data lakes, which is basically a data storage system. And then there, you know, it become immutable. Basically, those, data the data won't be
changed and you get the provenance of the data. Provenance means you know where the data originated, because there's something called proof of origin, where we can prove, basically, a data was generated on the certain sensor, right? Let's say a camera. And then put into a data lake and then from there into a smart contract to automatically run the
traffic flow in that smart city. And so we are able to obviously prove where the data came from and that is and it has a high degree of certainty at the same time, which is relevant, you know, for AI and your robot fleets also, too.
[laughter]
Definitely. I think one of the early use cases that people expected a blockchain
was to be able to put data to verify it. But it seems like there were there were issues whether the blockchain is, you know, a huge amount of size, it's slow, there's a transaction cost to put each piece of data on the blockchain. At least this was the case with Bitcoin maybe Ethereum. Are all of those non-issues now that we have faster, newer blockchains?
I think they're still issues, you know, like most [clears throat] blockchains are just forks of other blockchains. And so the first one Bitcoin, you know, was made to transfer value from peer-to-peer and to have an accounting ledger for it, right? That's that's the basic premise of Bitcoin. There's some people who tried to save data on there, save some
pictures and stuff like that, but it's crazy expensive, you know, there's if you try to save an image, you know, on the Bitcoin ledger, you know, you spend millions of dollars. And even on Ethereum, you know, it's going to be $100,000, $200,000. Now it's getting cheaper, but it's still very expensive. Solana still in thousands of dollar range. And Solana is a super modern,
efficient blockchain, right? And the reason is they are made to transfer this token value, you know, and not really made for data. And actually only a blockchain It's blockchain really built for data from scratch. You know, we started our deep in network, more than eight years ago and we had a data company basically. And
then so over the last eight years we said, "Okay, someone else needs to build a layer one for data." But nobody has done it. And so we built so many components over over the years. And then we said, "Okay, we have all the components almost, you know, let's just put it together and put it into this blockchain so we can solve it for other people as well." And that's how the XY
layer one was born.
Mhm. Yeah, you know, it makes complete sense. Comparing transactions to data, you know, transaction, I'm sure the size of it is fairly small. Although, if you look at Bitcoin, there's still quite a fee to send a transaction. But data can be [clears throat] huge and now that everyone is, you know, going so hard on AI, some of these
big companies like Microsoft, they told their employees, you know, use it AI as much as you can and they're just like making so much data.
[clears throat]
How does XY XL1, the layer one, how does it handle, you know, huge quantities of data and are you looking for quality versus quantity or it doesn't matter?
Yeah, it's it's handled so that we have a
infrastructure called data lakes, you know, which lets you store data and then and then you hash it and put that hash onto the blockchain itself. And this way, you know, if someone tampered with data, it was our data to save a lot of data itself on the blockchain. You can do that, you know, our blockchain is very efficient, but it would be better if you have huge
quantities, you know, to put it in one of those data lakes. And [snorts] So, what was the second part of your question?
[clears throat]
well, I'm just worried about the quality of data.
Quality. Yeah, quality is im- important, you know, So, the more you can prove data is, the less you have issues with hallucinations or wrong output even from your regular
team, you know, which which even even on AI team, you know, you need good input data, you know? Otherwise, you write something which might be wrong, you know, like if you look at America, America seems to have two realities, you know, like there's one narrative, you know, on the left side and one narrative on the right side, all right? And it's sometimes
one wonders, you know, what happened there, you know, how how can that happen because the day- the data underlying data must be the same, you know, someone either switches the data in the narrative or there's like some fear-mongering going on, like what's happening really, you know, that we have these two streams of thinking and so
I think by looking at the underlying data, you can see, okay, do we actually really have a problem with, I don't know, let's say immigration, for example, or what does in- what the impact does immigration have on the economy, right? And can look at on the data and have reasonable conversations about that, you know, and then have try to have one
reality, you know, yeah, and but you know, data gets gets faked and spoofed somewhere, you know, for people's gain, you know, to win elections, for example. And we try to solve that. So, quality of data is hugely important in all and most aspects, but more important than other in others like like political reasons, you know, where where it's high
stakes, you know, our health care, you know, when automated smart city or automated robots in the future, right? You want to make sure, you know, that they run on good data. And then, you know, the quality is better and more important than than quantity.
Definitely. If we were to look at the major AI chatbots right now, OpenAI, Anthropic, Google,
[clears throat]
how easy would it be for them to start tying the data into a DePIN network? Would they have to start from scratch? Is that something they could do overnight? Would they be incentivized to do it?
Yeah, you know, it depends on the DePIN network, you know, there's there's a lot of them, you know, there's like temperature sensors, there's
Things which create Wi-Fi networks, cell networks, and there's others which which measure other things, and there's sound and mapping and lots of different DePINs. And they generate the data in a decentralized way. The cool thing there is you know, you're able to contribute data, and that data then you earn tokens for that, you know, so you earn money in
return for providing the data, so you become part of the economy. And some but sometimes it's difficult they're insular, you know, they're decentralized as a DePIN network, but they're difficult to plug into, you know? And so, you need to be either like a Solidity blockchain developer or you need to have some other tricks up your sleeve to tap into into that data.
And so, at XY O, we looked at that, you know, a few years ago, and we said, "Oh, we need to make it really easy for regular web two companies and others, you know, to just plug into our network." We built a lot of stuff on like JavaScript and other easily understandable languages use APIs so that you can plug into into our system because
a lot of blockchain projects do blockchain for the sake of blockchain and they say, "Okay, we are blockchain project, you know?" And we say, "No, no, no, no, we are a project which has blockchain and has tokens and everything and uses that type of infrastructure and technology, but actually we are data company. And for that data company and data product, you know, how can we
maximize our adjusted market, you know? How can we bring the most people in to contribute data and connect with it to take out the data. And that's how how we have built it. But in the future Anthropic and friends, you know, in our case they are already able to plug into it, you know, like the next hour if they want to. We built a cool tool, tell you about it in a
second, but also others will do the same, I think. And we just last week released our new XY O AI SDK, which is an incredible tool allowing you to build any product on top of the XY O Layer 1. You can easily connect into it like even if you're a byte coder, you know, like or a or like a you know, you're a DJ or
Genesis teacher, you know, it doesn't matter, you know, you have an idea, you know, you want to connect to this blockchain, collect connect otherwise with it, otherwise with the DePIN network or you want to have immutable data or you want to create a game with your friends and which connects into the real world, anything you can think of, right? You know, can
now build on the XY O Layer 1 blockchain and you can use the XY O AI SDK to just code it up. And so it also allows you to easily connect into that data stream. So, if you need some data from our DePIN network, you know, you're able to easily connect into it with the XY O AI SDK.
Wow, that's very cool. For YouTubers and podcasters, would they be able to
upload their content or just hash their content to verify the provenance of it?
Yes, absolutely. You can do that. And even cooler like you can build a product like you can build a clone of Ashton, for example, right? It's responsive fans. And based on the content and you can make sure that it's only your original content your fans knows only your
original content. You didn't add some other content to it or if Ashton virtual Ashton gives bad advice, then you know, they can see, okay, what is that based on? And see the underlying data.
That's interesting. You know, there's so many YouTube videos recently that are AI generated and like Google's barely puts any warning labels or notifications on
it. It's like everyone's listening to music nowadays and they don't even know that it's AI produced. They're like, "This is so good." And it's like once you've heard it a bunch, it's it all sounds pretty similar. But, I feel like that's a quality issue. May- It sounds good, but, you know, peop- people are misinformed that it's not human created. So, is that something
that XYO would be able to ensure that people can determine whether things are AI created or not?
yes, yeah, because we're able to show the provenance of it. So, like the creator will be able to identify themselves, you know, that they did it and they say played trumpet, for example, right? And people knows that this person is a trumpet player and it's validated and
verified, you know, maybe that's real trumpet. The piece I played, right? But the Google and friends, you know, they are incentivized to have more YouTube streams, for example, right? So, for them, they're not incentivized to figure out what is AI generated and what not. You know, they want to show more ads. They show more ads if you are more on YouTube, you know? And so, you have
to be entertained.
Yeah, it's a it's always a tough one. As long as people are watching, they don't really care if it's AI or not. And I know I think we need some AI accountability. I saw the recent partnership with Theta, which is an OG project as well. They were trying to build like the decentralized YouTube and many other things back in the day and they're still
around. It's it's great to see. Can you talk about the AI partnership with XYO and Theta and how that ties into Data Lake and the SDK?
Yeah. We are very excited about it. Theta is a great project. You know, they started in 2019. So, just a year after us, you know, they have some great partnerships with Houston Rockets Olympic Games SA and then some other
sports team, you know, they for example use their [clears throat] AI agents for customer services and interactions there. And they built but at the core, they have a compute infrastructure network. They mostly excuse for AI compute and they are also creating AI agents. So, a [snorts] problem for AI compute
is super hot, right? So, all the those companies obviously are growing in value enormously, you know? If you look at Google and Amazon and everybody who builds compute there, you know, is has grown in value a lot over the last few years. And the same is true on the crypto side. And there Ether is Ether ICP and many others, you know. And so but what they lack is a true validation
of does an infrastructure actually exist, you know? Because a lot of projects are incentivized to have their token increase in value, so they might say, "Okay, you know, that they have maybe more nodes than they actually have." And what we are doing is verifying the AI compute infrastructure of Theta and they're also verifying some and we record
the answers of the AI agents and we can check that they're working as they're supposed to work. And you know, we record the answers and make it immutable, so you know we can it's a it's good for liability where you can say, "No, no, it gave the right advice." You know, here or it's you can check the quality of the answers and make sure that your agents run the way they're
supposed to run. And the cool thing is usually this project to build something like that usually would take months, you know? But it just took us a few days with the XY O AI SDK. And we are excited to build many many more of these kinds of partnerships over the next few weeks and months.
It's very nice. Well, congratulations on that. I do
appreciate Theta and NX and XYO. So, I'm excited to see that. I'm going to check back into Theta. It's been a it's been a short while since I looked into the details of their projects, so now that it's AI verified through XYO, I would love to jump back in. And I want to test out the SDK as well. You know, if a janitor can do it, I can definitely
toy with it and from the content side and the distribution and the AI virtual AI versions of myself, what's the best way to get started with that?
Yeah, it's probably go to extraordinal network website. They have a big big well, piece of our website is now about the extraordinal AI SDK. You can just you know, download it and connect there.
Sign [snorts] up for early access and then get started coding. And if you run into any problems, you know, just send an email to partnerships at extraordinal network, you know, and we can maybe partner with you or help you to get something there. Otherwise, if you're not quite ready yet, then, you know, maybe connect with us on X with official
XYO. Or sign up for our newsletter and, you know, we teach you over time to do it. But even if you're non-technical, it's super super easy to interact with the extraordinal AI SDK. It might be your first step into building something cool.
That's very cool. And with this just being released, is there a plan to focus on this heavily and do iterations or
you know, what else might be coming up in the road map that the team is focusing on?
Yeah, there's going to be iterations of this. It's going to be easier and easier, you know, like right now you need to have more prompts, you know, like give it more advice, you know, like how to build something. And but we [clears throat] help you because
you threw it right now. But in the future, you know, I'd love it to say, you know, build a clone of Ashton,
One prompt.
One prompt, you know, and just grabs all the data and does all the things, you know, and have it be very very very easy. So, we're working towards that. But, it's right now it's it's a good thing to start. And what's going to
happen there is we think over the next few years that millions of products are probably going to be built XYO AI SDK and that those products will connect with the XYO layer one blockchain as it is underlying infrastructure and data layer and so on. And so, it will increase transactions on the layer one by multitude and with that, you know,
like the XL1 token which runs the blockchain, you know, the gas usage and so on. So, it's it's an exciting time, you know, for the XYO XL1.
Definitely. And when you use the SDK to develop, do you does it require XL1 as a gas or can you actually earn it? Is it Is it a circular economy?
Right now, we give you a free XL1 to connect and build it build some things,
you know, on the testnet at least and then for later just reach out, you know, if you're it's it's very, very, very cheap. You know, except if you have something which is like tens of millions of users, big, you know, then you know, then it gets more expensive, but we are happy to talk with you about, you know, grants and stuff like that. You know, it's early we wanted to get
started and have fun, you know, I mean and nothing should impede that, but the transaction the transaction costs on the XYO layer one are are very low for any any data related tasks.
Definitely. No, that's great to hear and the experimenters and the tinkerers usually reap the benefits of early experimentation. So, I'm definitely going to check it out. I'll
leave a link to the main XYO network site, the socials you mentioned, if there's any other information on the SDK and for the experimenters, I'll put all that in the show notes below. Congratulations on the partnership for Theta the new data lakes everything else as AI grows. It's hard to keep up. You guys are doing a great job and I'm looking forward to
seeing more decentralized networks and deep in networks tied into AI as AI just bleeds into every corner of every industry. We need to have data provenance and security and blockchain networks there too.
I agree. Thank you. Actually, this is great.
Related coverage
More interviews
How Robinhood built its own blockchain to control compliance and feesSep 29, 2026
Geo Protocol founder Yaniv Tal launches Geo Debates, its first productSep 28, 2026
How Brave is expanding beyond ad blocking with Drew PotterSep 27, 2026
Why Kevin Carter thinks China is winning the AI raceSep 25, 2026
Brave's plan to bring self-custody payments to 120 million usersSep 18, 2026
Arie Trouw on why AI agents now need proof of actionSep 14, 2026