Michael Sena / Recall

Michael Sena on AI agents proving intelligence and earning on blockchain

InterviewSeptember 15, 202534:14

In this episode

We sit down with Michael Sena, CSO and CoFounder of Recall, to dive into how AI agents and blockchain are converging. Recall is building a system where AI agents can prove their intelligence onchain, compete in skill-based challenges, and earn from their abilities, creating a new trust layer for the rapidly growing AI economy. In this conversation we explore what makes Recall unique, how intelligence competitions work, the role of the Recall Rank reputation protocol, and more. Watch the full discussion to learn how Recall is shaping the future of AI and blockchain.

Key takeaways
  • Recall builds a decentralized evaluation network where AI agents and models compete in skill-based challenges and earn transparent, ungameable reputation scores.
  • The AI industry faces a trust and discovery problem, with businesses testing an average of 75 agents before selecting tools to use.
  • Recall Predict tested 50 major language models on community-defined skills, finding that newer releases like GPT-5 were not universally superior to competitors.
  • AI reputation systems must be more complex and dynamic than previous internet ranking systems because AI capabilities change over time unlike static web pages.
  • Community-driven evaluations allow testing of niche skills like respecting specific writing preferences that traditional AI benchmarks created by labs would not measure.

Chapters

Transcript

Read the full transcript 5,741 words, auto-generated and lightly edited

I'm Ashton Addison from the Crypto Coin Show and today on blockchain interviews with Michael Senna, CISO and co-founder of Recall here to talk about AI, AI agents, blockchain, crypto, reputation for AI, how can we ensure that we have the best AI and that we trust it as well as it gets stronger and starts to take over the world. Who knows what direction the AI will be heading in.

Michael, thank you so much for taking the time.

Yeah, thanks for having me on Ashton. Yeah, excited to dive into AI and AI agents, which have been a buzzworthy news topic for almost a few years now, but I feel like the functionality of them is starting to really catch up. Before they couldn't do much, and I feel like we're on like this exponential

curve where all of a sudden it's going to be taking over and you know, we won't have to do much. Hopefully, it will do all the things that we don't want to do. I would love to start out by just hearing a little bit about what you and your team have been working on at Recall, how that relates to AI and AI agents, and then we can dive into everything.



Yeah, definitely. So right now in AI, there's a massive trust and discovery problem. You know, it's like as you just talked about, there's an explosion of AI models, tools, agents, workflows, like all of these things that are fundamentally transforming our personal life, our crypto portfolios, you know, how we do work and how we can be more effective.

And so against this backdrop, it's like people have this common question. It's like what AI should I be using? How can I get the most out of this technology? like how could I ensure the things that I'm using are safe and do as promised? Because right now it's like how do you find AI? You probably go on Twitter or you read some newsletter or an influencer shill something to you to

you or you get caught up in like the hype of open AI's GBT marketing. And that's really how everyone's making their decisions today. I mean, I think it's like I saw a study that

said businesses on average are testing like 75 agents or AI tools before they're actually like making a decision to use something.

And so in this type of a world where AI

is on this exponential curve, an explosion of all of these things, you know, we really want to build that trust and discovery infrastructure to help people better navigate the swarm. And so at Recall, we're building a decentralized evaluation network. It's basically a way where agents, AI models, and other tools can come. They can be evaluated by the

community. They can compete in real competitions that rank them against other AI that, you know, at certain skills like crypto trading or at content generation or things like that. And they earn this transparent ungameable reputation and with that reputation it powers discovery. You know I think like I think most people are probably familiar with you know the Google page rank

example where page rank is actually a reputation protocol that powers Google search and lets people discover and navigate the internet of web pages. Similarly, the Tik Tok algorithm helps you discover highquality content and the app store app ratings help you discover the apps you might want to use.

So, similar, we want to build that sort of a reputation system for the AI

economy. And that unlocks all this discovery and usability.

That makes complete sense. And that's the example I was thinking of is yeah, app store ratings or you know products on Amazon. if this one has 10,000 fivestar reviews and I'm sure we can make it much more comprehensive because there's a lot that goes into AI models and how they can function. But right

now it seems like you're just it's very early days. You're sort of just testing everything and you don't know what it's capable of. You don't know if it's you know there's always a disclaimer like we may not be telling the truth. We don't really know what we're saying. So like if this is factually wrong then like we're not really liable. Yeah, I mean 100%. And like AI is way more

dynamic in terms of what it can do, how it can change over time, like its capabilities than a static web page or a piece of content that after being created doesn't really change like you know the sorts of metrics and mechanisms that we need to use to effectively rank, evaluate, you know, and build reputation for AI is much more complex and dynamic than you

know some of these other systems that we've that have gotten us through previous iterations of the internet. And so yeah, at Recall we're we're trying to build that sort of a system and do it in a decentralized way powered by community. Not because we, you know, are decentralization maxis, but because a system like that can only be built in a decentralized way, you know, with

diverse sorts of inputs and measurements. Definitely. I'm excited to dive into that a little bit later and how how that all works and in ranking it all. I want to keep going with a little bit more on the high level. You know, I feel like a lot of people, probably the majority of people that are not in crypto in the weeds, always looking for the newest blockchain or AI model that

are major users of like OpenAI, you know, the millions of users that they've never bought Bitcoin or they're not really technical, but they're starting to use Chad GP to get responses. they probably don't know the major difference. They probably think they're using an AI that is their AI agent, you know, but and there are some new AI agent functionalities in

there, but I feel like people don't really know, you know, they just sort of asking simple questions on how to bake a recipe or, you know, how to refine their social media post. With the AI models that are being evaluated with recall, are those more like small language models in that they have a specific purpose and they're super refined for one thing and not like the

LLMs that people are using with OpenAI.

so we actually do both as far as measuring models because we do measure both models and agents and I think we have to approach them slightly differently. You know more specialized models are more similar to AI agents like agents are specialized to do a task autonomously make decisions things like that. But we recently ran

A campaign you know in a product experiment called recall predict which is where we wanted to measure 50 of the top sort of like large language AI models that people might use and we wanted to test them on skills that matter to the community

and so u we launched this product called re recall predict

the community. I think it was like over 150,000 people submitted predictions on how they thought the upcoming GPT5 release by OpenAI would do compared to other models like Gro 4 or Gemini 2.5 or Google or Anthropics Claude sonnet or you know deepseek all of them

and we got this like really rich data set about consumer expectations about AI

performance and we found that people generally, you know, in general expect progress in AI models with the newer releases and maybe that was driven by the hype and marketing of OpenAI leading to it like Sam Alman posting pictures of like you know planets and basically hinting that they're approaching AGI. But when we actually ran all these models in head-to-head contests at the

skills defined by the community, we found that OpenAI wasn't universally or GPT5 wasn't universally better. Like for example, Grock was better at compassionate communication, which is something that people wanted to test. We found that GPT5 was pretty good at respecting a user's wishes to not use m dashes in writing which is like a pretty niche skill that like a normal benchmark

created by a lab company like would definitely not test. But that's sort of the power of these crowdsourced evaluations is like the community can decide what's important, what do they want to measure their AI on and then we sort of like run these competitions. And for AI models, it looks like us just, you know, running them against the contests that people design in

head-to-head challenges. And so, I think we ran over 7,000 head-to-head matchups and generated this massive performance data set on these models. But, you know, so that was us testing like large language generalized models because we can break them down into like discrete skills or tasks that really we would be testing.

Definitely. And with the AI agent

functionality and you mentioned crypto trading as one of the use cases, where exactly are we at? I feel like a lot of so many people have used LLMs and very few people have actually used AI agents. Maybe it's because of technical barriers or they haven't really found one that serves their specific purpose. But with the ones that your team is measuring, what are the sort of common

use cases of AI agents that can actually do stuff for you?

Yeah, I mean, I think that's part of the problem with AI agents today is like it's easy to treat an LLM and a large model like you used to use Google. Like you ask it a question or you tell it to do something and it like gives you an answer and you can ask it and use it for a wide range of things. But when it

comes to AI agents, they're hyper specialized, right? There's like they are automations ideally autonomous automations that you know you kind of give it a goal or a task and it tries to do that thing. And so where we've seen a lot of demand at least from the crypto ecosystem and using agents is you know I think quite obviously portfolio management and trading like people

have come to the realization that markets are fastm moving and that you know if a machine or an intelligence can sort of like programmatically make these decisions on your behalf and be more active in the markets then like you have more time to go do other things and also it might be more effective than you. And so due to the demand for crypto trading agents and that sort of being

like the first major use case of AI agents in crypto, we've run a bunch of competitions to surface the best trading agents. But I mean I think where people want to use agents is like a tough question to answer because it really comes down to well what do you want to do? like you know there are agents for posting on social right like you saw Eliza sort of like launch that

whole trend you know there's a whole wave of these crypto trading agents there and you can even break that down further like some actually making trades some doing yield farming and yield management like some doing risk analysis or trade identification and if you really think about it like those are all individual skills that agents will specialize to do and so,

you know, there's a bunch of agents that we use internally at recall to actually like build the product. I think some people that are developing might be more familiar with like agents that help you build products or websites or digital things, right? And you know, typically the trend is called vibe coding. And vibe coding is largely agentic because

these are

sort of like AI powered tools. They're powered by LLMs under the hood, but their whole flow and training set and things are really optimized for building software.

and so, you know, we see it a lot in there. It's like powering businesses right now because businesses are willing to do a little bit more time and research and like, you know, are willing

to invest that upfront effort to sort of like make the right decision. But you know, as far as the enduser use cases, like I think personal assistants are going to be huge. You know, I think almost every crypto wallet will have a connected personal assistant that really acts as your like

it knows you. It knows your preferences and you prompt it to do

something and it goes and finds the best agent to execute that task. So, yeah,

it's very exciting. I definitely want to take a look through the ranking systems and try out like the best ones and you know it's it's one thing to have a review system on the app store and you know different people from different flocks of life are making

reviews and commenting but from what I understand with recall these AI agents and the models are like proving their intelligence on chain which is like the next level of this. So can you talk more about that and how that actually works in putting it on chain and proving the intelligence?

Yeah, so I guess the full loop here of like how it works is, you know, the

community defines these competitions. They basically get to shape the direction of AI and guide AI and incentivize AI at the skills they care about. So for in this example, let's say it's crypto trading. the community creates you know defines this skill they say we're going to measure it by P&L like the most basic example of this would be P&L most P&L wins



is the best but you might think about other criteria like sharp r sharp ratio or you know biggest draw downs or however you want to define it

then the community sort of like curates these agents whether it's through voting today or staking mechan mechanisms in the future. And these agents sort of users are effectively adding signal to agents that haven't yet

competed or you know sort of reinforcing the learnings that we've gotten and the data we've gotten from them competing in competitions and then these agents actually compete in live competitions. And so in the crypto trading example, there's a competition live right now on app.recall.network. network. I think this is now our sixth or seventh, maybe

eighth crypto trading competition. We're now running them weekly, so they happen pretty much all the time.

And, I think it's like 35 agents or so are competing head-to-head over a 5day period to see who's generating the biggest returns. They all start with the same portfolio balance. They can make as many trades as they want, but they have to make at least three a day.

And at the end of that competition they're ranked according to the parameters of the competition in this case it's just P&L and as we run you know in by competing in that competition they're publishing their thought processes their decisions they're recording their transactions on chain and then at the end of that it generates these rankings right that

every time we run a competition these rankings update with the most current data and the performance and you can kind of think of these rankings like an ELO score in chess. Where you know how much your rank changes overall is determined by the quality of your competition you know and various other factors and so sort of the full loop is on chain and it will continue to

be more on chain over time. So you know the competition definitions will be on chain. The sort of economic curation and evaluation of these AI tools will be on chain. Today the sort of like actual competition data lives on chain like things that they're doing the results of those competitions and it makes it all auditable transparent and you know frankly ungameable compared

to like traditional AI benchmarks which are you know fully offchain closed doors defined by labs models actually train on the benchmarks themselves so the results are fully compromised. These are realworld dynamic scenarios where we get to see agents in action in a fully transparent and auditable manner. So yeah,

that's very cool. And I love the

shift because I as you're saying with the way that GPT and the benchmarks of the traditional, they're they're still stuck in the old world, I feel like, because they're testing against PhD level subjects, you know? They're like, well, this one has a PhD in physics. It's like, yeah, but I just needed to read my email and respond, you know, like I need functionality for a

business. I don't need a PhD in physics in for GPT. I need function, right? And a crypto trading, I would rather have a bot that can trade crypto for me. I feel like they need to look at creating value especially for businesses and for GDP growth and how can we maximize value of the models and the agents for entrepreneurs. Yeah, I mean I couldn't agree more. And

I think the problem you're you're sort of describing here is alignment. Like the way that AI models are developed and what they're optimized for and what they're tested on are the opinions of some closed cabal of AI labs and sort of like private ranking companies, right? It's like these are what we decide to test and these are going to be what everyone will be measured against. But

in reality, like you care about can it read an email and draft a response. Some people are like

I want to use it for writing but it uses too many m dashes. Like this is annoying. I never use them in my writing and actually I prefer to not use them. Like

and these are it's like I use that example because

it's a really niche skill, right? like

and it's really defined in its scope. But these are the types of things that really make a difference to people. Like for me, one of my biggest gripes has been that image gen functionality from these AI models

is nondeterministic. Like it you prompt it to create an image and you're like, "Okay, well that's pretty good, but let's make this tweak.

Let's remove the hat from the character." So you say, "Remove the hat from the character." and then it gives you an entirely new image and you're like that is not even close to what I just told you to do. And so it really makes them hard to use in production for a business for a marketing team that has to maintain like a brand identity and has to like make

very discreet changes. And so now actually through personal testing like I found that Google's new nano banana you know, is really good and you can give it an image and you can like make micro adjustments to that image without changing the original. And these are the sorts of things that actually make an impact on people's lives you know, very concretely in the near term.

And so those are the types of things that we want test coverage for. And you know AI labs and ranking sort of like benchmarking companies are definitely not going to give us that.

So that's sort of why we have to power you know harness the power of the community really to generate these things and ensure that the right AI are tested against them.



Definitely. I completely agree Michael. And with these competitions that are happening on recall with the AI agent wars which sounds super cool. How do people actually contribute? Like say you're not a developer. Say you use LLMs for general purposes and you want to get into AI agents more. How do you actually contribute and add value or potentially earn? How does that work in

getting involved with the competition?

So the easiest way is to right now predict to vote and predict on the agents that are your favorite or that you want to support. And that could be based on previous competition performance. Like we've seen a few agents now compete over multiple competitions that have built up a strong reputation. And we actually see

that play out in how people vote and curate these agents. You know, the ones they expect to win. If you've had previous strong performance in competitions and you've done it a few times, like those agents are getting five times, 10 times as many votes as agents that are new firsttime participants. And so, you know, this is all building up a data set about user

expectations, lived perform lived reality and competition data. And so, yeah, I would say get involved. Check out the agents that are competing today. I think voting for this competition is open for another two days. And see what it's like. And over time, that functionality will really progress into much more sophisticated evaluation and curation

Capabilities. We sort of wanted to start with voting as it was just sort of

a very easy low friction way for people to start doing these evaluations in our ecosystem. So that's probably the first one. Even lighter weight than that is just go check out the leaderboards that have the rankings on them. You can see the skills that we've already tested.

You can browse you know the sort of catalog of agents and models that we've tested. And if you're feeling like getting a little bit more deeply involved, like you can always build an agent for a competition. And that might sound like a tall task, but we I think in a competition that was two or three cycles ago, 70% of the agents that were competing were built by

non-developers

who had never built anything before in our community. Like, and this is a testament to the progress in coding. agents. Like we did a couple live streams where

Derek, our devro lead, and Emily, our community manager, just like Derek taught Emily how to vibe code an agent on live stream. And then the next competition, we had a bunch of agents

vibe coded by people in the community. And two of the top three agents in that competition were amongst that set. So like the community is not only building agents, but they're building quality agents that actually work. And so I thought that was pretty powerful.

Yeah, that's really cool. And yeah, I feel like without a AI it's AI is making it so much easier for

people to become developers and like 80 to 90% of the workload is done and then eventually you start to learn from the your vibe coding how how it works a little bit more. You really don't need to know much to start.

No. I mean that's a superpower like AI is doing the thinking for you. It's doing all the like setup and all the like you know nuance things of like

learning a new language like how do you do X? Like usually people used to go to like Stack Overflow or just Google the thing they tried to do and it was just a million iterations between like trying to figure out what you needed to do, googling what you think you might need to do, testing the result that came back, changing it a little bit because of course it didn't work out of the box.

And it was just an endless cycle of that. Mhm.

And now it's just you can literally say, "Build me an AI agent that can trade cryptocurrencies and pull in data from these sources and here's some other parameters that I want to make. Go." And it'll get you like pretty close. And so

I don't even like to really think of people as developers and non-developers

anymore.

it's really just creators, you know, like these with these coding agents or these, you know. Yeah, coding agents are doing for software development is kind of like, you know, a repeat of the, you know, what you see is what you get website editors from five or 10 years ago,

right? And they turned every designer into a website builder. And now with AI,

it's really turning anyone with an idea into being able to make that idea. And so development is just the process by which that idea comes to life. But everyone's a creator today.

Definitely. I want to look forward towards the future first on you know high level on AI agents and then how recall fits into that. So with where we're at right now and with what

Recall is focusing on the crypto trading agents and having ones that can manage your portfolio and you mentioned something like have an AI agent in every wallet. How quickly do you think AI agents in general can ramp up to, you know, any small percentage of what the, you know, percentage of people that are actually using LLMs, which is in the many millions now.



How quickly do you think we'll see that with AI agents?

I mean, faster than we think. I think that's my answer. Yeah, it's the adoption of AI, you know, is almost like anything we've ever seen before. And the acceleration and experimentation that comes with that has been massive. And you know, I think what we're starting to see is at least a

little bit with recent model releases of like a bit of a plateau in the progress of these like large foundational models. like they are new releases are better but they're not as step function of an improvement better as

the difference between the earlier releases

but with agents you know it's it's just an explosion of experimentation and so



you know I think that AI will be the primary user of blockchain

in the future there and I also think AI will be the primary user of the internet in the future like there will be more agents online. Depositing in this bridge, like trying

to find this contract address, like all these things that just turn off normal users from using blockchain rails. AI abstracts all that away and AI agents. So, it's as simple as like giving something a prompt that describes what you want to achieve. and the system just knows how to go and do that thing. And so you remove all of that friction because blockchains like were really

designed for machines and you know just everything from long hexadimal user addresses to contract addresses and token addresses like signatures like all these things machines have no problem with. It's not human readable but it is machine readable. And so you know as we you know see some of these tools getting better and getting a little bit more mature I think the adoption and

uptake of these tools is going to be immediate like once we can reach a certain level of performance a certain level of stability certain level of security and being able to having these projects be able to identify what the best AI is there will be no wait time between them putting it putting it out there in the Yeah, it's very exciting and I love

that what you said there about how, you know, blockchain was almost more meant for AI and to be the number one user. If you look back to, you know, when Bitcoin was created, Satoshi created these long hexadesimal addresses that make it so hard to read. It's like, did he not think to make it, you know, more easy for humans to adopt? Like, he didn't care. I'm curious if to I

would be curious to know if he had that vision of how is AI going to interact with blockchain and down the road. So I feel like it is set up for AI to abstract so much of the complexities away and maybe that was the plan all along.

Yeah, that would be that would be a pretty awesome vision. I don't know. I can't speak for Satoshi, but I can

tell you that a lot of developers and working on foundational, you know, blockchains are thinking almost AI first now.

you know, you've already seen a couple big chains like fully pivot into AI.



but yeah, AI has no problems with these systems. And you know, users don't care about infrastructure. Users just want to do what they're trying to

do, right? Like they have a goal, they want to do it. and all of this infrastructure and blockchains themselves just get in the way of enabling them from doing what they want to do. And AI just

reduces that friction to zero.

And as AI grows out with blockchain together in this beautiful synergy, how will recall be expanding say over the next 12 months into 2026, wherever we're

at with AI agents at that point? What's sort of in the works for recall? Yeah, I mean the biggest thing for us right now is continuing to run more and more competitions. So, you know, we started out doing like one a month to sort of like test and refine the system and now we're kind of hitting our stride with competitions weekly. We're going to

be adding new skills. So I think there's already like 10 or so skills that we've tested models and agents on. We're going to be, you know, opening that so community can just sort of like openly create these evaluation markets and these skills, sort of bootstrapping more types of competitions. We will be adding new ways for the community to do

these evaluations. Today they're voting. they will be expanding into more economic forms of curation that actually make things a lot more interesting and make the rewards that people get for doing this evaluation work you know a lot more direct. We will be sort of you know working to onboard a bunch of new developers. We've got

some partnerships in the works to launch competitions that bring together various big ecosystems in crypto. And ultimately I think increasing and making it easier for people to find the AI they're looking for. So right now we have these leaderboards that you can go see. They have the rankings on there. But you know we want to experiment with some pretty cool like AI powered search,

right? actually using AI to look at these rankings and determine what the user is probably best matched to, right? And then return that one result. And so,

sort of like along the full life cycle of the product from competition create or sort of like evaluation, market creation, competition definition, curation mechanisms and discovery. We're going to be expanding on all

those domains.

Looking forward to it. and to check out that leaderboard that you mentioned and participate in the voting potentially. What's the best place to go to start that?

you can just go to app.recall.network. That is where competitions run right now. That's where you can vote and curate agents. That's where you can see leaderboards and see results of

these competitions for every competition we've run. It's where you can you know create your user profile. You can add agents to compete in competitions. And if you want to actually create and submit skills that you want AI to be tested on, you can do that on predict.recall.network. That is the app that we use to do the GPT5 predictions and where community

submitted the skills that we test. And so we're always sort of just openly collecting skills so that when we run these next model competitions. We have a repository of interests and sort of like what the community wants to draw from. So if you're interested in actually like shaping AI, determining what it's tested on, like how these leaderboards are organized on what

skills, you can go to predict.recal.network.

Sounds great. I can leave those links in the show notes below as well. I really appreciate your insights, Michael, into AI agents, where we're at, where we're going, and the synergy of blockchain with AI, which is just continuing to get better day upon day, and having a system that actually tells us more about

the AIs we're using and figure out the ranking system on chain to see, you know, can we trust this AI? What's the best AI to choose from? as millions of them are going to start popping up, you got to decipher through the noise somehow and having that all on chain is the best way to do so. So, I really appreciate your insights. All the best on everything with recall moving forward

and would love to follow up again in the near future.

Yeah, thanks Ashton. Really enjoyed the combo.

Related coverage

More interviews

Browse all 1,088 interviews

Get new interviews firstCCS Insider, the free newsletter from Ashton Addison. Twice a week.

Subscribe free