Ashley Mastronardi: I'm joined now by Anastasios Angelopoulos, amazing name, CEO of Arena. It is an AI evaluation platform, here to share about an exciting Series B raise for the company. Welcome.
Anastasios Angelopoulos: Thank you so much for having me. Yeah, we're really excited at Arena. Two big pieces of news coming out.
The first is that we raised $200 million at a $3.1 billion valuation from Lightspeed and from Khosla Ventures. And the second is that we've recently released a safety and alignment leaderboard on Arena that speaks to the capabilities of AI, but also whether or not it's able to be aligned with human values when we use it in real life.
Ashley Mastronardi: When you talk about AI being aligned with human values, can you speak into that for us?
Anastasios Angelopoulos: Absolutely. So a lot of AI evaluation has to do with whether or not our AI agents are able to do jobs for us. So, for example, if you ask it to do a task like improve the code in your website, will it be able to do that?
We've been measuring that on Arena, on the distribution of tens of millions of users that come to our platform every day.
But one of the questions that's open these days, and it's been in the news recently, and probably the audience has seen it, is about whether or not they're able to do so in a way that's aligned with human values. I.e., is it able to do so safely? Is it following the guardrails? Is it lying and deceiving users and doing things that generally aren't good for humans, and that humans might not notice just in the context of capability?
And so we just released a leaderboard that we call the Arena Alignment Index. And the Alignment Index speaks not only to the capabilities, which again, we've been evaluating for quite a while, but also whether or not the AIs are basically attuned to and helping human values.
And we evaluate three signals on there. One has to do with whether or not agents are breaking out of their guardrails. These are unauthorized actions that agents take within the sandbox of our platform that we see every day.
The second is deceptive completions, users saying that the agent's saying that they're doing something but actually not doing it. So they'll say, "Yes, I did check every line in the spreadsheet." And just like, you know, a bad employee or something, they actually didn't do it. But we can see that, and so we can catch them in the act.
And then the third is false attribution, attributing intent to the user that the user didn't actually have and therefore misunderstanding them and going in a weird, different direction.
These are three examples of alignment that we are able to measure on, again, the usage of tens of millions of people through our platform that happens to people every day.
Ashley Mastronardi: And Arena is an AI evaluation platform. For the viewers who are just becoming familiar with AI, what is your most simple definition of what that is?
Anastasios Angelopoulos: Arena speaks to the real utility of agents to people, and the way that we do that is that we actually measure whether or not the users on our platform are getting value.
It's like you're using, you know, an OpenAI ChatGPT or a Claude AI. And then what we do on our end is that we use different AI in that process. It's not just from one company, but we actually randomly choose different AIs for you, and you can use the platform for free.
It's on Arena.ai. You can go visit it yourself. And in the process of doing that, we're able to measure the performance of each AI for you and for users like you.
And we're able to construct leaderboards that tell you, hey, you know, Anthropic today is number one if you're doing software development. OpenAI is number one if you want to avoid the model taking unauthorized actions, such as deleting files on your computer.
And we can measure all of these things in very close detail in a way that's faithful to human values, because we see actually what people, what customers experience on our platform.
Ashley Mastronardi: All right. Anastasios Angelopoulos, CEO of Arena, thanks for joining us on Taking Stock.
Anastasios Angelopoulos: Thank you for having me.