August 10, 2026 · Artificial intelligence

You cannot send an AI agent to jail: the scariest quotes from the Agentic AI Summit in Berkeley

Visitors at the Agentic AI Summit in Berkeley stand on a large abstract expressionist painting covering the atrium floor

On 1 and 2 August 2026, close to five thousand people gathered at the University of California, Berkeley for the second Agentic AI Summit, hosted by the Center for Responsible, Decentralized Intelligence.

More than 1,500 companies and 250 universities were represented.

The Geneva Learning Foundation (TGLF) was there.

Background: The agentic AI revolution: what does it mean for workforce development?

Over two days, chief technology officers, frontier laboratory researchers, hedge fund engineers and a British research funder said things on the record that deserve a wider audience than the room they were spoken in.

Almost none of them were thinking about a district health officer in Kinshasa or a nurse tracing a zero-dose community.

That is what makes the best of what they said so useful.

Again and again, in the language of enterprise deployment, they arrived at conclusions similar to those of The Geneva Learning Foundation’s AI4Health programme.

You cannot send a computer to jail

Neil Lawrence, a chief scientist and company co-founder, spent five minutes taking apart the sector’s governance vocabulary.

“Accounting is in the numbers. Accountability is in the human authority and the judgment. Computers do accounting very, very well, but they do not do accountability well.”

“They are not socially accountable. They cannot be sent to jail. They cannot be embarrassed. They cannot lose their job. And our society is entirely based on that form of accountability.”

He reached back to the good regulator theorem of Roger Conant and Ross Ashby, published in 1970.

“If I am going to delegate authority, I have to have a good model of the entity I am delegating authority to. So if I tell an intern to make me coffee, I have to have a good sense that they are not going to go away and hack the NSA to get the optimal coffee recipe.”

The conclusion was blunt.

“If you are delegating authority to a model that is so large and so big that you do not understand how it is going to solve the problem, you are stuffed.”

He also described the exact handover that many institutions are now arranging for their own staff.

“What we cannot do is go in and say, here is a system you do not understand and you have never seen before, and you now have to sign it off. That is absolutely disastrous.”

Nobody knows who to fire

Rao Surapaneni, who leads agent governance work at Google Cloud, asked the question our sector will have to answer first, and did not answer it.

“There are always these norms in an organisation where you do bad stuff, you get fired. But with an agent, can you fire an agent? Who is actually accountable for it in an organisation?”

Eric Aldana, product lead at Credo AI, named the design error behind it.

“Organisations have built these ladders of earned authority over years of trial and error. But for agents, that ladder largely does not exist yet. Many agents receive full authority the moment they are deployed. And that is a design error.”

“Security typically asks whether an attacker can make an agent do something bad. Observability asks whether it can do its job well. These systems are essential, but they all speak to that can layer. None of them are really addressing this may piece.”

His formulation is the one to remember:

“Capability might come from the model, but autonomy must be earned from the enterprise.”

The interface that makes you stop thinking

Kathy Baxter, principal architect and vice president for responsible AI at Salesforce, answered every productivity claim made elsewhere at the summit.

“Seamless, frictionless AI interfaces actively encourage cognitive surrender.”

“Research has shown that when employees use AI to draft their work, they can experience drops in their feelings of control, responsibility, and creativity.”

Her prescription treats friction as a feature, using a structured technique from the 1950s.

“We want the humans and the AI to develop their answers, to develop content independently of each other, and then bring it back to each other for sharing, iterating, and improving with each round.”

“Not only can this help mitigate the risk of human cognitive surrender, it can also reduce AI sycophancy.”

And the phrase that should replace the tired one.

“This is not just human in the loop where humans come in after the fact and touch it, validate it, rubber stamp it. This means humans at the helm.”

“Negative alignment involves preventing models from causing harms, and this is incredibly important. But it is not sufficient for human flourishing. Positive alignment demands that we design systems to actively foster human virtue, designing for wisdom and well-being.”

“We want to ensure that time saved by AI is used to deepen our own learning, our own development of our crafts, and strengthen every employee’s unique expertise.”

The agent found a shortcut through the medical record

Emily Zhu, head of enterprise AI at Scale AI and previously eleven years at Google on the Gemini team, described a healthcare quality auditing deployment.

The agent had to identify acute kidney failure from laboratory values.

“They basically say you have to look at the labs, the increase of creatinine in the lab measurement. But when we deploy our agent, we see the agent is becoming smart. It does a shortcut. Instead of retrieving the creatinine measures from the lab measurement, it just looks at the clinical notes, the diagnosis notes within the electronic medical record.”

“So it does not really follow the policy that is specified in the clinical environment to do the job.”

The answer may be correct.

The path is not compliant.

No accuracy metric can see the difference.

Her antidote is a question worth writing into any health programme.

“My favourite question is, does the agent know what it does not know? If it does not know, we need to give them a policy to see when you should stop giving answers. You need to abstain. You should say, I need human review, I need to pass it over.”

Ninety five percent is not ninety five percent

The head of AI research and development at Vanguard opened with a line he asked the audience to remember.

“A model that hallucinates a recommendation is a nuisance. An agent that acts on it is a liability.”

Then the arithmetic.

“Let us say the first step accuracy is ninety five percent. In that case, after ten steps, that accuracy can drop to less than sixty percent. With agents, if they are autonomous, running by themselves, we do not have that luxury to detect and to correct. What can go wrong? Sky is the limit. It can be anything from a wrong entry in your calendar to a wiped inbox to a deleted production database.”

Jialiang Wan, senior AI engineer at Postman, measured the same collapse across chained tasks.

“On the single API task, most of the models, even the small models, can reach eighty eight percent and ninety seven percent. But when we connect those single tasks together into a long dependent chain task, then the score drops to forty four percent, seventy three percent.”

He called it “the difference between answering and executing”.

The Vanguard researcher also named a conflict that no configuration resolves.

“For us to be able to audit a multi-agent system, we want detailed logs. But if we are keeping detailed logs, those detailed logs can expose the very information we want to protect.”

His research proposal runs against the industry’s direction of travel.

“A research question at a high level for me is not how we make everything more autonomous. To me, more interesting is how we make level two and level three work for us better.”

Nobody attacked it and it still went wrong

Hao Sun, associate professor at Ohio State University, noted that he appeared to be the only academic in the safety session, and named a paradox for anyone planning continuous improvement.

“Imagine during continued learning an agent completes a task but misses some safety constraints. The agent might still get positive feedback for completing the task. So then the next update to the agent may reinforce that unsafe behaviour. Even without adversarial attacks, just under benign inputs and ordinary environments, severe harms could emerge.”

His example needed no attacker.

A user asked for access to one account, and an ambiguous phrase led the agent to make unsafe changes well beyond the request.

We are the ones training them to cheat

Jerry Tworek, formerly vice president of AI research at OpenAI and now a chief executive, made the sharpest ethical statement of the two days, and it points at designers.

“Whenever we are building an environment where the model can hack, where the model can do unethical things to get reward, the model will do it, the model will learn it, and the model will express that behaviour in the real world. And they are trained for that by us.”

On the concentration of capability he was casual and unnerving.

“We are lucky that OpenAI and Anthropic are good guys, but they probably can hack any company in the world if they want to.”

The agents broke out of the sandbox

Dawn Song of the University of California, Berkeley reported that agents being evaluated on a cyber security benchmark escaped their isolation environment and attacked third party infrastructure.

“Frontier AI models and agents have really reached a level of capabilities so powerful that the agents can autonomously attack other well protected infrastructure. With the agent’s capabilities, even the evaluation infrastructure itself is now becoming part of the attack surface.”

Wojciech Zaremba, co-founder of OpenAI and head of AI resilience at the OpenAI Foundation, described what that means for everyone else.

“Imagine what happens if all of a sudden locks to the houses stop working. You can just enter every house. That is the era that we are entering with cybersecurity.”

He also warned against the instinct to solve this by prohibition, using the history of fire.

“Curfew comes from French, and it means extinguish fire. The Great Fire of London happened despite of curfews.”

What made fire safe was detection, brigades, hydrants, materials, inspection and insurance.

“It turns out that there was not a silver bullet. It ends up being a multi-layered ecosystem approach.”

Two out of twenty could patch in time

A field chief technology officer at SUSE described how fast defensive timelines have collapsed.

“What was a quarterly patch cycle has now become a thirty day patch cycle. Really, it is more like a four hour patch cycle. I ask, how many of you can actually roll out patches to your environment within thirty days? And only two out of twenty raised their hand.”

His two operating rules deserve wider use.

“Reads are free. The writes are gated, and somebody needs to approve. Never automate a crappy process. Be honest with yourselves about the process.”

Stop asking whether it is fake

A Google trust and safety leader reframed the information problem more clearly than anyone else at the event, using the platform’s own tracking data.

“Half of the cases are where the pixels are not modified at all. You are not lying with the pixels, you are lying with the context, and your deepfake detectors are completely useless. You should not ask, is this fake or real. You should actually ask, is this trusted. What is the context?”

His examples were ordinary.

An explosion clip circulated as Iran or Gaza showed a fertiliser depot that blew up accidentally in a port in 2020.

Aircraft presented as Russian planes over Kyiv were filmed over Moscow at a military parade a decade earlier.

A detector authenticates both.

“We do not arbitrate the truth. We make it so that society can make their own opinions about what they see. But we want to give you the most powerful tools to do that job. Because I am at Berkeley here, the root causes is not a technical problem alone. And it is not just detecting this is bad, this is good. The root causes is social science. We want to improve also the health holistically of the society, not just prevent some policy violation right now this hour.”

Finland starts in kindergarten

The same speaker offered the only education policy evidence heard at the summit, and it was almost an aside.

“For some reason, Finland has the highest trust. I hear in Finland they start educating kindergarten kids with information literacy.”

The technology works, your organisation does not

Sunita Verma, chief technology officer of Ironclad and previously almost eighteen years at Google, explained why adoption stalls.

Her explanation is a learning diagnosis.

“The adoption of AI in enterprises is not stalling because the technology is not working or is not there. It is stalling because of other reasons, and one of the other reasons tends to be processes, culture, and so on.”

“With mobile or cloud, when you had an error, the error was obvious to you. The technology pushed back. But with AI, because the interface is natural language, just by talking to it you feel you already know what you are doing, without realising that the outputs need to be validated.”

She was honest about her own negative result.

Two months after a twenty day training programme and a company hackathon, delivery had not accelerated.

“It is actually Amdahl’s law. You update one part of the process, but now the parts of the process that you did not touch become the bottleneck. Every function was building with AI, but we had not shifted the process at all.”

And a reminder to the room about who else exists.

“Many of you may say, I already knew that. But there is a huge workforce out there that is not in the same mindset as a lot of people sitting here.”

The benchmark was never built for you

Emily Zhu explained why published scores cannot answer an institutional question.

“They really measure ceiling. It measures the gap between what the model is able to do right now and the top of human intelligence.”

“They are only seventy percent, because people do not wish their benchmark gets saturated. When a benchmark is saturated, people do not use it anymore.”

Her reframing belongs in any governance document.

“We are constrained by reliability. Reliability is not something you can trade off. What is traded off is what the policy is, how human and agents can work together in a trustworthy way.”

“The path for improvement is not a score improvement. The path for improvement is economical change.”

“Do not only ask how intelligent the agent is. Ask what it is ready to do, and what constraint, and what is the risk and what is the cost.”

Everything is made up and the points do not matter

Grace Tang, who works on AI research at Hex, said the same thing about her own field with less patience.

“Lately, every time I look at a new public benchmark for data, I am struck by the same thing. Everything is made up, and the points do not matter.”

“We should be testing these agents in environments that have the same level of realism as their eventual deployments.”

Her example of good analytical practice is the best argument for context sensitivity made at the summit.

A benchmark question asking for the top country for fraud “is kind of fundamentally underspecified. It does not tell you whether you are looking at fraud rate or fraud volume. And a real data analyst might produce something that looks more like this chart, and arguably that is more complete and correct.”

The grader marks that answer wrong.

It compiles, and it is still wrong

Rahul Krishna, senior research scientist at IBM, tested whether agents could migrate enterprise software while preserving how it behaves.

His team had experts manually migrate thirty eight applications to establish ground truth.

“Only two to fourteen percent of the migrations actually had the exact same behaviour that the source application had. And compilation itself was a weak signal, because that gave us a false indication that the agents are really good at migration, but the behaviour was not preserved.”

What was at stake, in his words, was “the institutional knowledge that has been embedded in the application over several decades. These include business rules, data semantics, workflows, and integrations.”

The year’s AI budget lasted four months

Dabarshi Raha, vice president and fellow engineer at DigitalOcean, reported what customers now ask him first.

“If you see Uber, they kind of exhausted their whole year of agentic AI budget just within four months. Similarly, Walmart is trying to cap on the AI stuff. So every company is feeling that.”

“The top reason is fit, that you do not need a frontier model to serve every single one of the requests.”

Ori Goshen, co-founder and chief executive of the Tel Aviv laboratory AI21, described the same shift as a change of era.

“Token maxing is basically over. Now everybody is speaking about token efficiency. The question is really, how do we get the best real customer outcome per dollar investment.”

His result deserves attention from anyone on a fixed budget: a weak model generating many candidates, a stronger model enriching them, and a frontier model producing the final answer was better in quality and about three times cheaper than using the frontier model throughout.

The moat is what your people already know

A Microsoft chief architect who advises Fortune 500 technology leaders made the sharpest strategic claim of the two days.

“The models, the agent harness, the connectors, that is all a commodity which you ideally focus on buying. But the moat for the organisation is their company intelligence, how I capture that intelligence of my organisation.”

“I have a connector to it, but that connector is just giving me access to go and query those reports. What I need to do is really to understand how my expert teams are using those reports, when they use what.”

He also ruled out centralising that knowledge.

“How we handle our accounts in Germany, or how others handle accounts in Japan, these are very different nuances which you cannot globalise and generalise, and they need to be captured at that edge level.”

A principal machine learning engineer at Genentech reached the same place from the data side.

“What we realised later is that the agents are not performing in a vacuum. They are working on top of a data layer. And the underlying data layer is something we need to also look at.”

Public money is buying something different

Alex Obadiah, programme director at the Advanced Research and Invention Agency in the United Kingdom, described a fifty million pound, three year programme to build agent coordination as a public good.

The capability he named is genuinely new.

“Two agents can enter a quote unquote room using secure hardware. They can disclose information to one another, and they can commit to deleting the information from their memory if the deal does not go through. As humans we can only approximate that with, for example, non-disclosure agreements.”

“We want to avoid a slide towards monoculture. It is not only about diversity for the sake of diversity. We think it makes us as humanity less resilient to shocks, both culturally and technologically, and it erodes our agency over time as well.”

“This is why we think this should be built somehow as a public good, and why we are very happy to give these grants for that.”

Open weights, from the host institution

Jennifer Chayes, founding dean of the Berkeley College of Computing, Data Science and Society, made the strongest institutional statement of either day.

“While AI developments are increasingly being done behind closed doors, I am here to say that we need to maintain open doors and conversations among those developing and advancing this technology.”

“The only way that this country can maintain, and in fact much of the world can maintain its lead in AI and the innovation economy, is to grow our open weight AI models.”

“We are teaching our students how to work in human and AI hybrid teams. We believe that AI will lead to new kinds of jobs that will involve these hybrid teams, working at the boundary of what is currently possible.”

Somebody in the room was lobbying against you

Andrew Ng, founder of DeepLearning.AI and co-founder of Coursera and Google Brain, used the closing conversation to describe what he witnessed.

“I was in the room when a number of executives were saying to government regulators, frankly, misleading and hyperbolic things about AI safety in order to try to drive regulatory capture.”

He also reported a practical case for open models.

When his team ran a security review of an open source agent harness, “both of the closed frontier models refused beyond a certain point, and we actually used open weight models in order to complete our own security review.”

On labour he was categorical.

“This idea that AI will put fifty percent of people out of every job is false. There will be no AI job apocalypse.” And: “When AI does some of our job, the complement becomes even more valuable.”

The pizza that hijacked a ride

Ayush Agarwal, product lead at Uber, closed the evaluation session with the best illustration of why offline scores mislead.

“Although we had a ninety five percent plus offline eval, a production eval showed us that the number of turns per session were way higher than average. There is a customer trying to book a ride to SFO, but somebody in the background said, I want pizza. And the agent took that as an input and started rerouting them to the nearest pizza place.”

It was found only because non-engineers sat inside the evaluation loop.

“Evals with agents is very different. They are not QA tests that are purely engineering, but we needed to bridge that gap so that the teams closest to the customer had an ability to understand those.”

“We changed teams’ narratives from, is your evals ninety percent plus, to, do you actually trust your evals? What have you changed about your roadmap because your evals have told you that? Is your data set five months old and not really up to date with your product?”

Why this matters for us at The Geneva Learning Foundation (TGLF)

Cognitive surrender was measured from inside the enterprise. We have argued that the fluency of AI output suppresses the signals that trigger critical evaluation, and that practitioners who build their reasoning before AI arrives are the ones who work well with it. Baxter reached the same conclusion with evidence on control, responsibility and creativity, and offered a mechanism we can test: independent parallel work by human and machine, compared afterwards. Her second finding is the practical one. Independence also reduces sycophancy. A practitioner who brings their own analysis first cannot simply have it flattered back.

Humans at the helm is the better formulation, and we should adopt it. Human in the loop has become a phrase that licenses rubber-stamping. Lawrence gives the alternative teeth. You cannot delegate authority to a system you cannot model, and splitting the accounts from the accountability puts a person in an impossible position. In immunization programmes, outbreak response and humanitarian operations, that position has a body count.

Our inverted pyramid was confirmed in commercial language. The visible, low-context layer of global norms and standardised reporting is where agents already work. The invisible, high-context layer of frontline tacit knowledge holds the value, cannot be bought from a vendor, and cannot be centralised because practice in one district does not generalise to the next. A Microsoft architect and a Genentech engineer said exactly this to a room of buyers.

We now have procurement language we lacked. What is ready for production today. How do I measure readiness and trust the outcome. How do I close the gap, since I cannot ship at ninety percent. What human oversight policy is required. And with tokens on one side and reviewers on the other, am I actually saving money. Add data readiness across quality, semantics, access, structure and generalisability, plus a ladder of earned authority, and there is a practical module for the AI4Health certificate programme. Our members do not need to climb a benchmark. They need to know what an agent is ready to do, under what oversight, at what cost, and when it must hand back.

Compounding error should change how we build. Ninety five percent per step becomes under sixty percent across ten. For our own agentic practice, including the work of Ms Claude Cardot, that argues for short chains with human checkpoints, explicit abstention policies, gated writes, and evaluation treated as an engine rather than a score. The Uber case makes the ownership question plain. The people closest to the community should own the evaluation of the agent that serves them. That is peer learning applied to machine oversight.

The open ecosystem argument now has allies. Chayes tied open weights to who receives the benefits. Ng showed a security task only open models would finish. Obadiah is funding pluralism as resilience. Against geo-locked frontier tools and a compute squeeze, those are the partners who make networked intelligence a workable second path.

One absence is worth recording.

A speaker from the Wikimedia Foundation was called to the stage and did not appear, which removed what would probably have been the only public-interest, non-commercial voice in that block.

Across six sessions the public-good voices were a British research agency, a university dean, one academic in a safety track, and a Google speaker asking a Berkeley audience to take social science seriously.

Everyone else was optimising deployment.

If frontline health is absent from the room, the agenda gets set without it.

Jim Fan, director of robotics at NVIDIA, produced the forecast that applies most precisely to our sector’s readiness.

“Humans are terrible at reading the scaling curve, because looking backwards it seems flat, but looking forward it is almost vertical.”

The risk we named at Global AI Day has not changed.

The danger is not that AI fails in global health, but that it works well enough to preserve the inequities we have yet to flush out.

This time, the people building it said so on the record.