All posts
    Sibyl Insight

    Education After “Intelligence” Becomes a Commodity

    SibylVcSibylVcJuly 21, 2026

    As builders and parents, we spend a lot of time thinking about what makes software a good, enduring solution, while also pondering how best to educate our kids. Thinking through what makes an enduring technical solution naturally made us ponder what makes a resilient human in the age of AI.

    Machines can now do a whole layer of what we used to call intelligence: recall facts, run the calculation, produce a fluent answer. That layer is getting cheaper every month. So the real question for education is not how to make humans better at competing on that layer. It is what humans should learn to do once that layer is essentially free and universal.

    What the Game of Go showed

    When AlphaGo beat Lee Sedol in 2016 at the ancient Chinese game of Go, it looked like the old story: machine beats human, proof that the machine is better. But what happened next was more interesting. Human play got better.

    A study of more than 5.8 million professional moves from 1950 to 2021 found that, after superhuman AI arrived, professional players began making more original and more accurate decisions. They tried sacrifices and shapes they might once have dismissed because the machine had explored ground they had not. The emblem of that was “Move 37,” a choice so alien that commentators initially assumed AlphaGo had made a mistake. It had not. The move was brilliant, and it sat in a corner of the game that centuries of human players had never seriously considered.

    AlphaGo did not invent a new game. It discovered parts of an existing one that humans had not reached. It compressed into a few years exploration that might otherwise have taken generations or, in the case of Go, perhaps centuries.

    Recent developments in mathematics push this point beyond games. In May 2026, OpenAI reported that an internal model had autonomously disproved a longstanding conjecture in discrete geometry, with the proof subsequently checked by external mathematicians. On August 1, it published ten further advances in mathematics and theoretical computer science, each resolving or making substantial progress on a longstanding open problem and accompanied by a formal Lean certificate.

    These are not merely faster repetitions of work humans already knew how to do. They push beyond previously known answers and, in some cases, use connections mathematicians had not anticipated. But they still begin inside a mathematical world humans have built: a posed question, agreed definitions, standards of proof, and a community capable of determining why the result matters. Like AlphaGo, the model moved the frontier inside an inherited frame, even when it moved that frontier beyond any known human answer.

    In a sense, that is what every useful technology has done throughout history. Writing did not invent memory; it extended it. The calculator did not invent arithmetic; it performed it faster. AI applies the same pattern to an expanding range of intellectual work. The machine is neither simply a teacher nor an enemy. It is an accelerant. It covers and explores ground faster than we can, which should free us to concentrate on the places where human judgment remains most consequential.

    The scoreboard

    AlphaGo never invented Go. Humans built the board, wrote the rules, and decided what winning meant. The machine explored, brilliantly, inside a game it was handed. Even AlphaGo Zero, the later version trained without human game data, was still standing on the deepest human input of all: the definition of the game and its rules.

    A scoreboard exists only once someone has decided what counts as a point. Inside a clear scoreboard, a machine can become nearly untouchable because it can play itself millions of times and measure the result.

    Modern models can also help devise metrics, propose hypotheses, decompose problems, and suggest alternative frames. The recent mathematical results show that they can produce discoveries humans have not reached. The narrower and more durable point is that there is no neutral scoreboard for what deserves to be optimized.

    A metric becomes consequential because a person, institution, or society decides to use it. The highest leverage human work is often deciding what should count, noticing when a proxy has drifted away from its purpose, and accepting responsibility for what follows.

    It is the difference between mastering a Newtonian model and recognizing the boundaries of its domain. Einstein did not simply decide, in isolation, to replace Newtonian physics. Experimental anomalies, earlier scientific work, and existing mathematical developments constrained the discovery. The important educational point is not that a machine could never contribute to such a shift. It is that progress required recognizing that a powerful model was incomplete, not merely applying it more efficiently.

    Blind spots make the same point from the other side. In research published in 2023, a team trained adversarial policies to probe KataGo, one of the strongest Go engines ever built. The system discovered a strategy that caused KataGo to make serious mistakes. Human experts could then implement the strategy without algorithmic assistance during play and consistently beat the otherwise superhuman engine.

    The adversarial system was not generally better at Go. It was good at one particular thing: finding a boundary at which KataGo’s otherwise extraordinary competence failed. The vulnerability had also gone unnoticed through ordinary training and self play because KataGo rarely encountered opponents that approached the game in this unusual way.

    Put the two examples together and you have the argument. Inside a well defined game, the machine can accelerate past us and discover things we have never seen. But framing the game, detecting where its assumptions fail, deciding whether its scoreboard reflects what matters, and taking responsibility for the consequences remain the highest leverage places for human judgment.

    Not because machines will never participate in those activities, but because somebody must ultimately choose and own the frame.

    Education’s job is to prepare children for that level rather than confining them to the rung on which machines are already becoming stronger.

    The floor above

    Much of what makes a person valuable in this world sits one level above simply producing an answer. Roughly five connected skills live there.

    Framing the problem

    Real problems do not arrive prepackaged. “Cut food waste in half” has to be broken into parts before anyone can act on it: what to measure, what is habit versus logistics, and what to estimate versus verify.

    Schools usually perform this step for the student, handing over a clean question immediately after teaching the method required to answer it.

    AI can already produce substructures for larger problems, sometimes impressively. But decomposition is not neutral. The way a problem is divided determines what becomes visible, what disappears, and which possible solutions are considered. Students therefore need practice not only solving a supplied question, but deciding what the question contains.

    Detecting where the normal answer stops working

    Consider anyone who has compared how they felt after an exam for which they were extremely well prepared with how they felt after one for which they were less prepared.

    After the exam they were extremely well prepared for, they may be able to replay their answers and identify exactly which questions they got wrong and why. After the exam they were less prepared for, they are much fuzzier about what might be right or wrong. They made enough decisions based on what seemed highly probable that it becomes difficult to predict their final score.

    When an AI system is wrong, it can resemble that moderately well prepared person, although not necessarily because it lacks information. It may know a great deal while still failing to recognise that a familiar rule is being applied to an unusual case. It can produce the average or highest probability answer with reasonable confidence even when that answer does not quite fit. Nothing about the wrong answer necessarily announces that it is wrong.

    The future of education therefore needs to teach around the edges. Whenever possible, topics should be explored through three kinds of cases:

    1. Where the rule works.
    2. Where it partly works.
    3. Where it fails.

    Today, schools mostly test whether students can apply a rule after being told which rule is relevant. As that task becomes cheaper, education should increasingly test whether students understand the rule’s domain: when to apply it, when to modify it, and when to reject the frame entirely.

    The goal is to develop students who notice when they have moved beyond the point at which the average or highest probability answer still applies.

    Judging answers instead of just producing them

    The machine gives you a fluent answer cheaply. The scarce skill is looking at several answers and knowing which is strongest, what is wrong with the others, and what additional evidence would raise your confidence.

    Defending work aloud is the classroom version of this. It exposes the moments when a confident answer and a correct one have quietly separated.

    Knowing which fact matters

    More information is rarely an advantage by itself; machines can process more of it than any individual person. The advantage is knowing which piece of information changes the answer.

    Give students more material than they need and ask them to identify the fact that actually decides the case.

    Owning the call

    Someone still has to decide who receives a scarce resource, whose harm counts against whose benefit, and how to defend the decision when it goes wrong.

    A company can assign part of this process to a review board. It cannot make accountability disappear. Models can advise, compare options, and predict outcomes, but people and institutions still choose whether to act on those recommendations and bear the consequences.

    This belongs inside every subject, not in a separate ethics class, because responsibility becomes real only when something is actually at stake.

    These skills are connected, although they do not always operate in a strict sequence. Framing determines which answers are relevant. Boundary detection reveals when routine application is unsafe. Judgment identifies the evidence that matters, and accountability forces the entire chain to be defensible.

    Schools do not need to teach them as a rigid ladder. They do need to practise them together.

    But the foundation still matters

    None of this means dropping arithmetic, grammar, history, or chemistry. It means reconsidering which work students must perform unaided, which work they may delegate, and what they must still understand.

    Think of structural engineering. An engineer no longer has to calculate every load by hand or draft every drawing from scratch. Software can model forces, test designs, surface conflicts, and revise plans faster than paper ever could.

    But the engineer still has to understand what the software is doing. They need to know whether its assumptions are appropriate, whether the model reflects reality, and where the structure might fail.

    That is the role AI is beginning to play in education. It can produce the answer, run the calculation, and generate the first version faster. But students still need to understand the work underneath, because they can direct a machine, challenge it, or catch its mistakes only if they understand the thing it is doing for them.

    The honest objection is: if machines may eventually build the higher floors too, why learn any of this?

    Because understanding is what allows you to climb and steer. The person who only presses the button is at the mercy of whoever, or whatever, understands the system behind it. Understanding the foundation remains nonnegotiable. Personally laying every brick does not.

    A classroom that spends all its time laying bricks produces students who can build one floor and stop. A classroom that covers the foundations while preserving time to ask “so what?” develops skills that will remain valuable for students growing up with AI: debating whether a decision was right, distinguishing a disagreement about facts from one about values, and judging a choice from the perspective of someone it affects differently.

    You reach those questions by getting through the foundations, not by skipping them.

    This isn’t a leap into the unknown

    In our view, the panic that AI forces us to reinvent school completely is overstated. Some systems we have studied and experienced have been building these upper floors for decades, long before anyone worried about a chatbot taking people’s livelihoods.

    France’s four hour baccalauréat philosophy examination asks students to identify and formulate a problem, reason through it rigorously, evaluate arguments, examine a thesis, and justify a conclusion. Its Grand Oral asks students to present and defend a prepared question clearly and convincingly before examiners.

    The International Baccalaureate’s Theory of Knowledge asks students to reflect on the nature of knowledge and how we know what we claim to know. Creativity, Activity, Service connects education to experience, reflection, initiative, collaboration, and projects beyond conventional academic work.

    None of these systems solves the whole problem, and access to their strongest forms remains uneven. But the blueprints already exist.

    Others have circled the same territory. In Assigning AI: Seven Approaches for Students, Ethan and Lilach Mollick propose classroom uses that keep students actively responsible for supervising, questioning, and assessing AI output rather than becoming passive recipients of it. The OECD Learning Compass 2030 identifies creating new value, reconciling tensions and dilemmas, and taking responsibility as transformative competencies, years before the current AI wave.

    The destination is therefore familiar. The more specific proposal is to make boundary detection central: teaching students to recognise when a system is exploring the right game, when a rule has moved outside its valid domain, and when a confident answer is optimizing the wrong scoreboard.

    The three case pattern, where the rule works, partly works, and fails, is one practical way to teach it.

    What you can do tomorrow

    Schools move slowly. Curricula take years to change, and retraining teachers can take even longer. Parents do not necessarily have to wait for the whole system to catch up. Ordinary moments at home provide a faster experimental space.

    Here is what that might look like. A seven year old, fresh from watching Mary Poppins, wonders whether holding an umbrella could make him fly, based on a true story. The adults immediately stop any real world attempt. But instead of ending the discussion with a lecture, they treat the question as an opportunity to explore how scientific questions are investigated safely.

    What evidence would help answer the question? What smaller observation or demonstration could reveal something useful? What would the evidence need to show before the child changed their belief? What is the smallest version of the question still worth asking?

    That becomes an exercise in framing the problem, identifying which detail changes the answer, and revising a belief rather than accepting one answer on faith. No new curriculum is required. It simply requires seeing a moment of intense curiosity as an opening for thought.

    The parental instinct is often, and rightly, to keep children safe. The missed opportunity is stopping the conversation once safety has been restored. Children are most willing to think deeply when the question is already theirs.

    Kids choose stranger questions than any curriculum would: the umbrella, the bug they will not put down, or the argument over whether a game’s rules are fair.

    Each one is a chance to practise the whole process in miniature:

    What is really being asked? Where does the obvious answer break? What would a better answer look like? What should we check? What did we learn? Who bears the consequences if we are wrong?

    The machines will keep taking the floor below, and they may eventually help build parts of the one above.

    Our job, as schools and as parents, is to make sure children understand the building well enough to keep climbing.

    We explore this same commoditization dynamic from a company-building angle in Building a Moat After “Intelligence” Becomes a Commodity.

    Updated on Aug 3, 2026 to incorporate recent OpenAI advances in mathematics.

    ##ai##artificialintelligence##education#Education