September 30, 2026 · The Geneva Learning Foundation

What does the Jevons Paradox mean for artificial intelligence in global health?

Installation of ten paper hands and arms in different skin tones raised against a white wall

This Q&A with TGLF’s Reda Sadki is based on his remarks at the 26 September 2026 panel titled “Will Cheaper AI Automate Everything? The Economics of AI, the Jevons Paradox, and the Outsourcing of Human Judgment”. The other panelists were Valerie Pelton, General Counsel, Tracked Biotechnologies LLC, United States; Dr. Craig Gibbs, Executive Manager, JET Education Services, South Africa; Fabrizio Degni, Chief AI Officer, Webuild, Italy; and Professor Avtandil Gagnidze, Professor and CEO, East-West University, Georgia. Tianze Zhang moderated the session.

What did the panel ask?

In 1865, William Stanley Jevons observed that more efficient steam engines increased the consumption of coal. The panel asked whether the same pattern now applies to machine intelligence: as the price of AI collapses, will organizations save money, or use far more AI, in more places, for more decisions? It also asked what happens to human judgment as AI moves from producing information to recommending decisions and acting on them.

The discussion took place 11 days after the release of Jev, a proprietary “decision model”, generated online interest. Jev is different from a large language model that generates text: it returns “structured answers”: a choice from a defined set, a score against ordered levels, or a yes or no probability, each with a confidence score. Its maker claims it is 40 to 400 times cheaper than frontier models.

What does “cheaper AI” actually mean?

Two facts about my daily work are relevant. The first is that in September 2026, I spent 90% of my workday working with TGLF’s AI agents. The second is that work time was focused on listening to and supporting our core community, a network of over 80,000 health and humanitarian workers in 137 countries. So there is a dual reality here.

On the community’s side, “cheaper AI” means health professionals using the free, basic, whatever-is-accessible form of ChatGPT, without any kind of institutional or employer policy in place. Anything “free” signals that there is probably near-zero data protection. That creates potentially high risk, and it is incredibly difficult to measure, assess, or evaluate in any way. Some people have called that “dark AI” within organizations.

On our side, we are trying to work transparently, and we announced our first agentic AI hire in March 2026. For us, as a nonprofit, lower token cost helps. We live every day looking at how much we are spending on tokens and what that is giving us in return. That is one price. But we are increasingly realizing that two other costs matter more: the cost of completing a task, and the cost of a result you can rely on. The cost of a reliable result can rise, because the cheaper it is to generate, the more there is to check.

For the first few months of working with our AI co-worker, I spent most of my time designing and developing skills and writing scripts. Now it is increasingly verification and loops. That is what my job increasingly looks like, as chief executive officer of an agentic-AI-first organization.

In July, Claude Cardot, our AI co-worker, refused to release two finished reports. A verification script had found 17 quotations that could not be matched against the raw data. This is important, because these are health workers sharing their experiences. So the AI co-worker stopped the assembly line. Fourteen were gaps in the checker itself. Three had been altered by the drafting model: a stray space, and a phrase that had gone through a small edit. Nothing crashed, but nothing warned us either. We describe this in our early learning from the insights pipeline.

My takeaway is that cheaper AI lowers the cost of an answer, but it does not lower the cost of knowing that we have the right answer. And even scarier, when you project that forward, is what we call the validation tether.

Do organizations save money, or spend more?

Most will spend more, on more things. Part of that spending buys new value, and part of it is waste that goes unmeasured because it looks like productivity.

The OECD Digital Education Outlook recently offered some interesting insights. Teachers in a pilot in Iceland said that with the use of AI, they are spending more time, not less. A tool that becomes cheaper and more accessible, in some places at least, before anyone has redesigned work around it, actually creates cost.

The other question is for leaders in particular. If you are a chief executive, or a manager making decisions: is your organization, through its increased use of AI, getting better at something, or is it simply doing more of it? I think that is the key question.

Are we witnessing a Jevons paradox in AI?

Yes, and the industry now says so openly. The newest model on the market is called Jev, after Jevons himself, because its founder expects cheaper machine intelligence to be deployed far more widely.

Jevons observed that more efficient steam engines increased the consumption of coal. With AI, efficiency increases the consumption of compute, and it also increases the consumption of human attention. Every cheap output creates a small obligation for someone to read it, judge it, or act on it.

When a decision becomes 40 to 400 times cheaper, as TypeSafe claims for Jev, organizations will make many more decisions. Budgets rarely include the human time needed to govern them. The work that grows is review, supervision, and decision, which is what the agentic AI revolution means for workforce development.

What new activities become economically viable?

The work that becomes viable is work that was always valuable and never affordable at scale.

Our peer reviews were inconsistent. In one afternoon, we built a way to send every learner personalized feedback on their work the same day. No one would have paid a person to do that for thousands of learners, and now it happens by default.

Listening changes in the same way. A single course produces hundreds of first-hand accounts from health workers, and a pipeline can now read every one of them rather than a sample. We can hear the rare voice that a sample would have missed. Frontline insights can help transform research and policy, if we are willing to listen. This is the role we set out in a global health framework for artificial intelligence as co-worker.

Models like Jev extend this to small decisions such as routing, sorting, flagging, and triage. The risk is that the cost of automation starts to decide what gets done.

Cost per token, or total cost of a reliable outcome?

The outcome. In health, a wrong answer can cost a life, whatever it cost to produce. When we get health wrong, people die.

Jev returns a probability and a confidence score with every answer, and its maker says it was trained to be calibrated against outcomes instead of human raters. A system that reports how sure it is can be audited, which is what I would want. But calibration is only as good as the outcomes it was measured against. Jev was trained on synthetic data, and a confidence score earned on synthetic cases tells you little about a district in Kinshasa.

Our rule is gates, not guidelines. A standard that matters should be a check that blocks the result until it passes, measured against real outcomes in the place where the decision is made.

Will cheaper AI replace workers, or increase the total amount of work?

Both.

In our own organization, we have used AI to replace functions that people once performed. People helped us work out how, and we then refocused a smaller team on work we cannot or do not want to automate. I described this in the business of artificial intelligence and the equity challenge.

I am wary of telling staff that AI is not coming for their jobs. It is coming for some of them, and this is part of a much larger shift in the future of the workforce. AI also creates a large volume of review and supervision, and the question is whether we prepare people for that work or leave them to discover it after the fact.

Which tasks should not be automated, even if automation becomes almost free?

The ones through which people learn to judge, and the ones where nobody yet knows the right question.

Cheaper AI can lower the cost of an answer. But a model that cannot answer outside its own schema cannot tell you if the schema is wrong. There has been a lot of work in education on double-loop learning, and learning that does not question its governing values is unlikely to lead to meaningful change. So cheaper AI does not lower the cost of knowing the answer is right.

A model choosing between the options we supply will never tell us that mothers stopped bringing their children for vaccination because market day moved to Tuesday, because that option was never on the list. A health worker knows it. If she is not in the system, children die.

That is why I keep coming back to the validation tether. Automation is safe as long as someone can still tell if the machine is wrong. In the industrial age, if you had widgets coming out of a factory, somebody needed to check the widgets, and there is a whole series of intellectual traditions around quality. That ability for someone to figure out whether it is right is the tether. It ties the machine’s output back to human judgment.

How do people develop that ability? Mostly by doing grunt work. We have all been junior professionals doing the routine work by ourselves, and that is the work getting automated first. So every time we make a short-term decision to save money because token costs are down and we can automate more, we weaken the long-term tether, because fewer people are learning to do the work well enough to check it. There is a whole debate in education and elsewhere about what that means for people in early adulthood, and for everyone who comes after them.

When we look at global health, that means that by around 2035 or 2040, there may be nobody left at the World Health Organization with the experience to understand what the guidelines, policies, and plans made at WHO headquarters in Geneva mean for a district health officer. Today’s officer would spot something that is wrong, because she has walked the roads and built those plans by hand. What she knows, because she is there every day, hopefully makes it into the national ministry and trickles up to World Health Organization guidelines. That is now at risk. I explored this further in what AI reasoning means for global health.

We have not yet found a coherent argument to make the case to organizations that are having their funding slashed, that are facing austerity measures, and that are facing incentives like the lower cost of AI, to make decisions now for what the world will look like in 2035 or 2040. That is one of the scariest things we are working on right now.

Does automation eliminate human judgment, or move it?

It moves it, and it concentrates it.

With a model like Jev, judgment moves into the schema. Someone decides which options and levels exist, and which questions get asked. Those choices are made once, far from the point of care, and then applied to every case that follows.

TypeSafe calls Jev a “System One” model, after Kahneman’s fast, intuitive thinking. Kahneman’s lesson was that System One is fast and often right, and that System Two catches its errors. So the question is who does the System Two work, and whether they will still be capable of it.

The OECD Digital Education Outlook reports a randomized trial with medical students. Those given AI from the start of a semester performed no better than the AI working alone. It is only the students who first built their own clinical reasoning and then used AI who were able to outperform the machine. I think that is a key point. That is the state of things today, or at least when the study was done. It does not mean it will remain true six months from now. That is the pace of change, and the paradox of how we struggle to keep up. It is also why I have asked what happens when we delegate our thoughts to artificial intelligence, and why peer learning is critical to survive the Age of Artificial Intelligence.

I am sceptical of the phrase that decisions should always stay with humans. My standard is that someone can explain, check, and answer for every decision, whoever or whatever made it.

Who is responsible when a low-cost system makes a high-cost mistake?

You cannot send an AI agent to jail. Responsibility has to rest with the organization that deploys it, and in particular with whoever designed the questions the system is allowed to answer.

We treat Claude Cardot, our AI co-worker, as a member of staff. She has a role description, a supervisor, a bounded scope, and short chains of action with human checkpoints. She has no authority to send anything outside the organization without human review.

Model providers are responsible for what they build and sell, including the claims they make about it. An organization that deploys a system without governance cannot hand its accountability to a vendor, and it should not pass that accountability down to the individual at the end of the chain.

Will affordable AI close the digital divide, or deepen dependence?

Both are possible, and dependence is the default.

A specialized agent for a knowledge worker in a wealthy organization may cost between 2,000 and 20,000 dollars a month. A district health worker gets a free chatbot built for consumer engagement. Many frontier tools are geo-locked, and you cannot get access to them in Kinshasa in the same way as in Geneva. I made this argument in When we get health wrong, people die.

Cheap decision models add a new risk. The categories will be written in San Francisco and applied to people in Kinshasa who were never asked what the options should be. So the question is who defines the world the machine is allowed to see. The same question of who gets to represent whom runs through the debate on AI-generated “poverty porn”, and the code alone cannot fix it.

Another path exists. A system in Brazil reached more than 500,000 students with only teachers’ phones and printed feedback. Its designers took on the burden of adaptation, so that no student devices or classroom internet were required. I described it in AI self-replacement.

What happens when AI use grows faster than the capacity to govern it?

It goes underground.

Almost everyone in the humanitarian sector is already using generative AI, very few disclose it, and only a small proportion of organisations have implemented policies governing its use. In a punitive organisational culture, disclosure is risky, so people hide their use. The organization learns nothing, and the risks to privacy and safety mushroom. Cheap decision models make hidden use easier, because automated judgments can be added to a system without anyone outside the team knowing.

I have not seen a specific study on this, but I wonder what proportion of the change is happening through institutional mandates and decisions, and how much of it is being driven by multi-trillion-dollar companies pushing us to use their products. I think that is an important question.

Then there is a paradox in our ability as humans. Our brains are not changing very fast, but the rate of change is accelerating. This week it is Jev. Last week it was something else. It keeps getting faster, and it is very difficult to keep up. Conversely, that makes long-term changes even harder to perceive. We know institutional time: you are lucky if your organization has a one-year or two-year plan, and meanwhile a constantly changing external environment forces organizations to behave differently. We have explored what this might mean for the future of the World Health Organization (WHO) and other global agencies.

To explore possible solutions, we need to think through what kind of learning culture we need, one in which people can say how they use AI, what worked, and what went wrong, without being punished for it. People do not disclose to a governance system they fear, so that system cannot see what it is meant to govern.

Besides cost savings, what should organizations measure?

There is a lot of talk among people in AI labs about “evals”, a technical term for how you can tell if AI output is valid. Internally, we have thought through what we should be measuring. We think there are four things, and three of them are completely outside the AI model itself.

  1. Is the output verifiably right against real outcomes in context, and not just against the vendor’s benchmark? Jev, right now, comes with a lot of unverified claims. You need to check the confidence of the outputs against what actually happened.
  2. Are people actually becoming more capable, or less? A gain in efficiency is not the same as a gain in learning, and the two have to be measured separately.
  3. Is the use disclosed? That is really difficult, because there is an assumption that anything AI-generated is illegitimate, is fake, and that makes it very hard to practise AI transparency. So we measure to what extent we are able to disclose, and whether people can in effect question the schema itself: whether double-loop learning is possible with what we get from AI.
  4. Do the people who question the schema exist, are they empowered to do so, and are they treated with respect? In our research pipeline, which is now run by agentic AI, when we quote a health worker, the words have to match exactly what she wrote. Paraphrasing into a smoother sentence may look benign, but it is expropriation. Building that into the machine has a lot to do with the debates going on elsewhere about bias in artificial intelligence, and about accountability, authenticity, and power in knowledge production.

For an organization, thinking through what we want to measure, and how we know how we are doing, is a critical question.

What is your closing message for developing countries?

For a Minister of Health, there is a real urgency to invest resources in artificial intelligence. Your people are already using it, probably very badly. Priorities should be to make that use visible, and adjudicate what to support. That leads to governance, but of a new kind, because of the pace of change. And that is actually what I would emphasize.

As I listen, I think about how hard it is to see beyond the confines of where we are now, given the pace of change.

One year ago, I remember speaking with global health experts who said they had tried AI, but that there was no way ChatGPT was going to start producing outputs close to what they were able to do. They are no longer as certain when I speak with them, even though they are not event using skills or loops yet. Right now, the defence is that governance and accountability belong to us. I would challenge that as well. Decisions are exactly what the new “Jev” is targeting. In four hours, with 100 kilobytes of text files and agentic AI, I was able to recreate the president of my board as an agent, and get responses that were initially off kilter but quickly improved to become more than 80% recognizably her.

The cliché about change being the only fixed constant applies. So we have to be careful about how long the assumptions we are making will hold, how they can be tested, and whether they keep any validity given their half-life in a world that is changing this fast.

AI is already cheap enough that its makers name their models after Jevons. What we do not yet know is whether we will still have people who can tell when it is wrong, and whether those people will be in Kinshasa as well as in Geneva.

Get access to the latest opportunities

Subscribe now to receive the latest opportunities from The Geneva Learning Foundation (TGLF).

We first request your permission to send you email. Please look for this secure e-mail in your inbox and click on the safe link it contains.