Intelligence Learns
How to Find Out.

Perception, action, and the making of evidence.

The world hands out outcomes, not lessons.

Turning one into the other is the work that belongs to the learner.

In this blog

In Where Do Good Vision Targets Come From?, I argued that the bottleneck can lie in the observations from which we build learning targets, not just in the loss used to learn from them. In Plato Is Not a Space, I argued that visual knowledge should be understood through the predictions about change it supports, not only through the similarity of internal spaces.

The two positions meet at a practical requirement. Testing a prediction requires observations that can support or challenge it. When the evidence at hand leaves competing predictions unresolved, another view or a revealing action can make their consequences distinguishable. Can a visual learner arrange those observations for itself?

This brings us back to an unresolved difficulty in the first blog. An agent might collect its own experience, but a kitchen does not come with a scoring function.[1] In a predefined task, we specify what to measure and how a result counts as success or error. Those choices are not automatically supplied for each new question about a kitchen.

It is tempting to make the problem disappear by saying that the world grades our predictions. Predict that a cup will tip, push it, and see what happens. But the world supplies an outcome, not an explanation of what the outcome establishes. The cup might leave the camera’s view. The push might differ from the one intended. A single outcome consistent with the prediction might still leave several explanations intact.

Turning what happens into something we can learn from is already work.

My position is that more of that work should belong to the learner. A general visual intelligence should not only improve its answers from the evidence it receives. It should learn how to obtain and interpret evidence for questions that matter and expand the range of questions it can investigate.

An outcome is not yet a learning signal

Suppose a robot wants to know whether two parts are rigidly attached. It pushes them, and both move. That observation does not settle the question. Contact and friction might have carried them together.

The robot could try to hold one part still while moving the other. It might first need a better view of the joint, a more reliable grip, or a way to distinguish the object’s motion from the camera’s. The relevant achievement is not merely executing an action. It is selecting an action and a viewpoint that make competing predictions distinguishable in the observations.

I use evidence here to mean an observation interpreted in relation to a question and to how it was obtained. It may change the support for a prediction without settling it. Obtaining useful evidence therefore requires linking a question to an action or observation and a measurement that bears on that question. None of these is guaranteed merely by putting a camera on a robot.

Even then, a gentle pull that produces no visible relative motion need not establish rigid attachment. Friction or motion below the camera’s resolution could produce the same observation. The result is evidence under particular conditions, not a complete explanation of the mechanism. Learning includes recognizing what remains unresolved.

From an outcome to something worth learningThree stages: an outcome, evidence for a question, and knowledge worth learning. Question and measurement connect the first two; purposes and future use connect the last two.01OutcomeWhat happened?Both parts moved.02EvidenceWhat does it establish?Separate competingexplanations.03Worth learningWhat does it enable?Use the distinctionelsewhere.Question + measurementPurposes + future use
From an outcome to something worth learningThree stages: an outcome, evidence for a question, and knowledge worth learning. Question and measurement connect the first two; purposes and future use connect the last two.01OutcomeWhat happened?Both parts moved.Question + measurement02EvidenceWhat does it establish?Separate competingexplanations.Purposes + future use03Worth learningWhat does it enable?Use the distinctionelsewhere.
Figure 1.From an outcome to something worth learning.A measurement bears on a question under stated conditions; it need not settle the question. Whether learning from it is worthwhile also depends on purposes and future use. The arrows identify work for the learner, not automatic implications.

This is also work that a dataset’s creators may already have done by choosing a view, recording an action, or retaining the frames that reveal a change. A recorded interaction can support learning without being repeated. A dataset of recorded observations is not outside the world; it preserves selected observations of it.

What agency adds is the ability to influence which observations the learner acquires next.

A new skill can be a new way of knowing

Once a learner can obtain evidence for one question, what does it retain that could help it investigate another? I care about visual knowledge through the predictions it supports about what will remain stable, what will change, and what would follow under different conditions.[2]

Such knowledge need not be verbal, or stored as a collection of explicit propositions. A broadly useful visual representation may support a new question through a task-specific readout learned after pretraining. The concern here is what happens when the available observation, or the learner’s way of using it, is insufficient.

Return to the robot and the two parts. Learning to hold one part steady might look like a manipulation skill. It can also make previously ambiguous motion informative. Learning to rotate an object can reveal surfaces that the original camera never saw. Learning to use a mirror can make a new line of sight available.

In each case, an ability changes not only what the agent can accomplish, but what the world can teach it.

A new view may resolve the current uncertainty without improving the system’s ability to learn in a new situation. For the process to become developmental, the system must retain something reusable, such as a relation it discovered, a better way of measuring, or a way to arrange a revealing situation. A motor skill opens access; learning from the resulting evidence is a further achievement.

That is the developmental loop I want to build. Knowledge enables action, action can make informative observations accessible, and interpreting them can support further knowledge. Crucially, what is retained should help the system obtain or interpret evidence in another situation, rather than just repeat the original task.

A skill can open a new route to evidenceA cycle links what the system knows, what it can do, what evidence it can reach, and what it can learn next. Its focus is learning new ways to acquire evidence, not repeating a fixed loop.What it knowsPredict stability and changeWhat it can doHold, turn, and inspectEvidence it can reachMake a hidden relationobservableWhat it can learn nextReuse knowledge fora different questionA developmental loopLearning changeswhat can be learned.
A skill can open a new route to evidenceA cycle links what the system knows, what it can do, what evidence it can reach, and what it can learn next. Its focus is learning new ways to acquire evidence, not repeating a fixed loop.What it knowsPredict stabilityand changeWhat it can doHold, turn,and inspectEvidence it can reachReveal a hiddenrelationshipWhat it can learn nextUse knowledge fora different questionLearning changes what can be learned.
Figure 2.A skill can open a new route to evidence.A skill can make an informative observation accessible. The loop develops only when interpreting that observation leaves knowledge or a way of investigating that can be reused. This is a developmental possibility, not a guarantee of unlimited growth.

The goal is not to challenge every prediction indiscriminately. It is to learn how to identify predictions that fail or remain underdetermined, including those the model initially made with confidence.

Side noteOne representation, many questions

A broadly reusable encoder and a system that actively seeks evidence are complementary. A representation may already contain what a new question requires; the next step can be a different readout, not another experiment.

Steerable Visual Representations provides a concrete example of conditioning a visual encoder on supplied text to emphasize different concepts in the same image.[3] Such conditioning changes how available information is used. A prompt can also supply prior knowledge or a task description, but steering alone is not a new measurement of the pictured scene. Changing the question, acquiring evidence, and updating knowledge are different operations.

Learning changes what can be learned.

Not everything checkable is worth learning

Giving the learner a role in choosing its questions also allows it to choose badly. It could become excellent at generating easy questions, predicting irrelevant regularities, and repeatedly confirming what it already knows. High accuracy on self-generated tasks would not establish that the system had learned anything useful beyond those tasks.

The issue is not peculiar to vision. In discussions of AI-assisted mathematics, formal verification is distinguished from checking whether the formal statement captures the intended claim; a correct result is also distinguished from an intellectually valuable one.[4] Applied here, these are two separate requirements. Evidence must bear on the question actually being asked. Answering that question must also be worth the learner’s effort.

To say why a distinction is worth learning, we first need to say what the distinction amounts to. Calling the parts “rigidly attached” should make a difference to predictions about how their relative position changes under an applied force. Merely assigning a different label would not supply that content.

This is where pragmatism enters the argument. In How to Make Our Ideas Clear, Charles S. Peirce proposes clarifying a concept through the practical consequences it could conceivably have.[5] I take this as a way to connect a representation to predictions that can be tested. What would be different, under what conditions, if the distinction were correct?

That is a claim about meaning, not yet about which questions deserve attention. A distinction can have clear consequences and still be irrelevant to the learner’s purposes. Whether it matters also depends on the system’s goals, available sensory and action capabilities, and resources, a relationship emphasized by the Umwelt Representation Hypothesis.[6] None of this makes convenience a test of truth. A prediction that seems useful can still be wrong.

Consider the two-part assembly again. Suppose a visible indicator stays on independently of the connection. The same recorded interaction lets the learner check a prediction about the light and a prediction about the parts’ relative motion. Only the latter helps distinguish the connection in this example. Checkability alone does not determine which target is worth learning.

Verifiable predictions can differ in learning valueTwo prediction targets use the same recorded interaction. Under the goal of learning how parts are connected, predicting a constant independent indicator is verifiable but uninformative about the connection. Predicting relative motion is also verifiable and may inform the connection. Transfer to another assembly remains an open question.GOAL: LEARN HOW THE PARTS ARE CONNECTEDThe same recorded interaction supports two prediction targets.TARGET AIndicator statet0t1t2Can the prediction be checked?Yes—compare it with the observations.What does it tell us about the connection?The light stays on independently.Its state does not distinguish the connection.TARGET BRelative motionABt0ABt1Can the prediction be checked?Yes—compare it with the observations.What does it tell us about the connection?Relative motion can constrain the connectionunder the observed conditions.Checkability alone does not determine learning value.Better decisions and transfer to new tasks are further claims to test.
Verifiable predictions can differ in learning valueTwo prediction targets use the same recorded interaction. Under the goal of learning how parts are connected, predicting a constant independent indicator is verifiable but uninformative about the connection. Predicting relative motion is also verifiable and may inform the connection. Transfer to another assembly remains an open question.Goal: learn how the parts are connected.Same interaction, two prediction targets.TARGET AIndicator statet0t1t2Can the prediction be checked?Yes—compare it with the observations.What does it tell us about the connection?The light stays on independently.Its state does not distinguishthe connection.TARGET BRelative motionABt0ABt1Can the prediction be checked?Yes—compare it with the observations.What does it tell us about the connection?Relative motion can constrainthe connection under theobserved conditions.Checkability alone does notdetermine learning value.Better decisions and transfer to new tasksare further claims to test.
Figure 3.Verifiable predictions can differ in learning value.Illustrative comparison, not measured model performance. The indicator is assumed independent of the connection. Both targets can be checked from the same interaction; their relevance differs under the stated goal. Colour could matter for another task. Reuse on unfamiliar assemblies is a possibility to test, not a result established by this example.

My additional step is to include future learning among the consequences that make knowledge valuable. Knowing how to hold a part can make a joint observable; understanding that joint may help the learner inspect an unfamiliar mechanism. The immediate answer is only part of the return. If a way of investigating transfers, its value can extend beyond the task that first made it useful.

This gives the developmental loop a purpose without reducing it to today’s reward. The aim is not to collect every possible distinction, but to acquire knowledge that improves decisions or opens worthwhile paths of inquiry, under the learner’s purposes and constraints.

Some of the value of knowledge lies in what it makes possible to learn next.

World models for exploration and training

The learner now needs to choose an action or viewpoint likely to provide useful evidence. It can learn how to make that choice directly from experience, or use predictions to compare candidate actions. The second route gives a world model a concrete role. Here, a world model is a learned predictive model of how the environment and its observations may change, including under candidate actions.

Considering possible consequences before acting can support action selection.[7] Here, the choice includes actions whose immediate purpose is to learn. The model need not already know the answer; it needs to help identify an observation for which different possible answers predict different outcomes.

For the two parts, it might predict one motion pattern if the joint is rigid and another if contact alone carries them together. That comparison could suggest holding one part while watching their relative motion. The result still has to be interpreted under the measurement limits discussed earlier. If the real mechanism was absent from the imagined alternatives, observation must be allowed to revise the alternatives too.

The same model can also generate simulated experience for training. Rather than only selecting the next action in the real environment, a controllable world model can generate variations of a situation in which the agent learns or evaluates its behavior. This is the promise of using world models as resources for training and evaluation.[8] Knowledge then becomes part of the infrastructure for acquiring more knowledge.

But learning from simulated experience and obtaining new evidence from the current environment serve different roles. Simulation can expose implications of the model’s assumptions and support learning behavior under those assumptions. Generating more samples from the same model does not establish that it accurately describes the scene outside it. Improvements within the simulation still need to transfer to the real environment.

Model predictions guide interaction; observations can update the modelA world model generates candidate outcomes and simulated experience. Proposed actions lead to interaction with the real environment. Observations return from that interaction and can update predictions. A boundary separates model-generated outcomes from independently acquired observations.INSIDE THE MODELPredict action outcomesCompare, simulate, train.REAL INTERACTIONObserve outcomesOBSMeasure, interpret, revise.Select an actionLearn from evidenceAn imagined consequence is not a new measurement of the world.
Model predictions guide interaction; observations can update the modelA world model generates candidate outcomes and simulated experience. Proposed actions lead to interaction with the real environment. Observations return from that interaction and can update predictions. A boundary separates model-generated outcomes from independently acquired observations.INSIDE THE MODELPredict action outcomesCompare, simulate, train.REAL INTERACTIONObserve outcomesOBSMeasure, interpret, revise.Selectan actionLearn fromevidenceAn imagined consequence is nota new measurement of the world.
Figure 4.Model predictions guide interaction; observations can update the model.A world model can compare candidate actions and generate simulated experience for training. A new observation can test a prediction only if it measures a relevant difference. Model-generated experience is not an independent measurement of the real environment.
Side noteThinking longer, seeing something new

If a maze is fully visible, more computation may reveal a route already determined by the evidence. If a door’s state is hidden and no available clue settles it, drawing a detailed future does not reveal its actual state. A prior may favor one possibility without providing a new observation of that door.

Generated alternatives can still identify a useful next observation. Their vividness is not the criterion; their contribution to finding out is. This distinction does not depend on a particular generation architecture.

This separates two timescales rather than two fixed modules. Within a single interaction, the learner uses perception to select an action. Over repeated interactions, learning new actions and interpreting their outcomes can expand the observations it can acquire and interpret.

Within a decision, perception may precede action. Across learning, action can change what perception becomes capable of.

Generality beyond pretraining

For generality, I want both the ability to apply pretrained knowledge to new tasks and the ability to adapt through new experience when that knowledge is insufficient. Generalization from existing knowledge and adaptation through new observations are distinct capabilities. A useful pretrained representation supports the first and provides a starting point for the second.

Still, suppose two situations produce the same available observations but require different answers. Those observations alone cannot identify which situation the learner faces. Prior knowledge can make one answer more plausible, and more computation can extract overlooked implications; neither is a new observation of the unresolved difference. This is a limit of the evidence for that question, not a universal ceiling on passive learning.

When informative observations are accessible, the next challenge is acquiring and interpreting them without requiring us to redesign the data collection and learning targets for every new task. That extends the goal beyond answering unfamiliar questions from pretrained knowledge. The learner must also adapt how it acquires knowledge when that knowledge is insufficient.

Our contribution to Visual General Intelligence: A White Paper describes visual general intelligence (VGI) in terms of acquiring usable world structure from partial visual experience and applying it to new tasks.[9] This blog extends that acquisition process beyond experiences already prepared for the learner, to conditions it can learn to arrange for itself.

Side noteLimits and lineage

This is a research stance, not a guarantee that every question becomes answerable. Sensors, resources, safety constraints, and the world impose limits. An autonomous learner should discover useful intermediate questions while remaining accountable to the purposes and constraints it was given, rather than replacing them with whatever is easiest to measure.

Curiosity-driven developmental learning already studies how acquired skills become stepping stones for further learning.[10] The perception–action perspective likewise places autonomous experience collection inside the learning problem.[11] The commitment here is to make obtaining and interpreting useful evidence a transferable capability.

The argument requires neither a single internal representation nor a rendered video before every action. A learned way of acquiring evidence need not be represented as an explicit procedure. What matters is whether it helps the system learn from another situation.

My bet is that the ability to find out should itself transfer.

Transfer would mean more than executing the same successful motion on a new object. A learner should be able to adapt a way of revealing, isolating, comparing, or testing so that an unfamiliar situation becomes informative. The new task need not have been anticipated when that way of investigating was learned.

I expect this reusable ability to become a major source of generality in sustained physical interaction. For systems already equipped with broad pretrained visual knowledge, the stronger expectation is that improving how they acquire missing evidence will matter more for unfamiliar physical situations than simply accumulating more comparable passive observations. That priority is a further claim to test; it does not follow merely from the fact that active observation can help.

If each new family of questions still requires us to invent a fresh dataset, specify a fresh target, and build a fresh way to interpret success, then we have retained a substantial part of the learning process outside the agent.

Conversely, if newly learned ways of obtaining evidence remain useful only for their original tasks, the developmental loop I am betting on is much weaker than I hope.

The goal is not a machine that has already seen everything. It is a machine that can use what it knows to find out what it does not, then turn what it finds out into further capacity to learn.

The bet

Generality in vision will come less from what a system has seen than from what it has learned how to find out.

References & further reading

  1. 01
    Where Do Good Vision Targets Come From?↩Kohsuke Ide

    The starting question: where learning targets come from, and the difficulty of interpreting feedback outside a prepared task.

  2. 02
    Plato Is Not a Space↩Kohsuke Ide

    Visual knowledge considered through the predictions and responses a representation supports, rather than a prescribed internal geometry.

  3. 03
    Steerable Visual Representations↩Jona Ruthardt, Manu Gaur, Deva Ramanan, Makarand Tapaswi & Yuki M. Asano · 2026

    A concrete example of conditioning visual computation on a supplied question or concept; distinct from obtaining new sensory evidence.

  4. 04
    Mathematical methods and human thought in the age of AI↩Tanya Klowden & Terence Tao · 2026

    §4.4 separates formal correctness from fidelity to an intended mathematical statement; §§4.6–4.7 separate correctness from broader understanding and value. The application to visual learning here is an analogy, not a claim proved in that paper.

  5. 05
    How to Make Our Ideas Clear↩Charles S. Peirce · 1878

    A method for clarifying conceptual meaning through conceivable practical consequences. The further claim that knowledge can be valuable for enabling future learning is the position developed in this blog, not a consequence established by the maxim.

  6. 06
    The Umwelt Representation Hypothesis: Rethinking Universality↩Victoria Bosch, Rowan Sommers, Adrien Doerig & Tim C. Kietzmann · 2026

    Box 1 and §2 connect useful distinctions and representational alignment to goals, sensing and action capabilities, and resource constraints.

  7. 07
    Coding Is All You Need? Why We Need a World Model!↩Hokin Deng

    Considering consequences before acting; the companion blog discusses deliberation during video generation.

    Schrödinger’s Cat in Video Models
  8. 08
    World Models at Third Dimension AI↩Christian Rupprecht · 2026

    World models as controllable resources for training and evaluation, including the difficulty of generalizing beyond their own training data. The blog uses that role, not a guarantee that simulated experience transfers to reality.

  9. 09
    Visual General Intelligence: A White Paper↩Hirokatsu Kataoka et al. · 2026

    §2.10: usable knowledge acquired from partial visual experience and applied to prediction, imagination, reconstruction, and new tasks.

  10. 10
    Intrinsically Motivated Goal Exploration Processes with Automatic Curriculum Learning↩Sébastien Forestier, Rémy Portelas, Yoan Mollard & Pierre-Yves Oudeyer · 2017 / JMLR 2022

    Autonomous goal exploration and skills that become stepping stones for further learning.

  11. 11
    The flavor of the bitter lesson for computer vision↩Vincent Sitzmann · 2026

    The perception–action loop and the problem of acquiring experience autonomously.

知能は、
調べ方を学ぶ。

知覚と行動、そして証拠を得る条件をつくること。

世界が返すのは帰結であって、教訓ではない。

その帰結を学びに変える仕事を、学習する側に引き受けさせたい。

このブログの内容

『Where Do Good Vision Targets Come From?』では、学習のボトルネックは損失関数だけでなく、学習標的のもとになる観測にもあり得ると論じた。『Plato Is Not a Space』では、視覚知識を、内部空間の類似性だけでなく、変化についてどんな予測を支えられるかを通して捉えたいと論じた。

この二つの立場を結ぶのは、予測を確かめるための条件である。予測を支持したり、疑ったりできる観測が必要になる。手元の証拠では相容れない予測が残るとき、別の視点や、違いが現れるような行動によって、その帰結を見分けられることがある。そのような観測を、視覚の学習系自身が用意できるだろうか。

ここで、最初の記事に残った難しさへ戻る。エージェントは自分で経験を集められるかもしれない。だが、キッチンには採点関数が備わっていない。[1]あらかじめ定義した課題では、何を測り、どの結果を成功や誤りと見なすかを私たちが決めている。キッチンについて新しい問いを持つたびに、その決め方まで自動的に与えられるわけではない。

世界が予測を採点してくれる、と言えば、問題は消えたように見える。カップが倒れると予測し、押して、何が起こるかを見ればよい。しかし、世界が返すのは帰結であって、その帰結から何が言えるかという説明ではない。カップはカメラの視野から外れるかもしれない。実際の押し方が、意図したものと違うかもしれない。一度、予測どおりの結果が得られても、複数の説明が残ることはある。

起きたことを、そこから学べるものへ変える。それ自体が、すでに一つの仕事である。

私の立場は、その仕事のより多くを、学習する側に担わせたいというものだ。汎用的な視覚知能は、受け取った証拠から答えを改善するだけでなく、重要な問いについて証拠を得て読み取る方法を学び、自分が調べられる問いの範囲を広げるべきだ。

帰結は、そのまま学習信号になるわけではない

ロボットが、二つの部品が固く接合されているかを知りたいとする。押してみると、両方が動いた。それだけでは、問いに決着はつかない。接触や摩擦によって、一緒に動いただけかもしれない。

片方を静止させ、もう片方を動かしてみることはできる。その前に、接合部がよく見える視点、より確実な把持、あるいは物体の動きとカメラの動きを区別する方法が必要かもしれない。ここで重要なのは、単に行動を実行することではない。競合する予測の違いが観測に現れるように、行動と視点を選ぶことだ。

ここでいう証拠とは、ある問いと、その観測が得られた条件に照らして解釈された観測である。問いに決着をつけなくても、予測をどこまで支持するかを変えることはある。有用な証拠を得るには、問い、行動や観測、そしてその問いに関係する測定を結びつけなければならない。ロボットにカメラを載せただけでは、そのどれも保証されない。

それでも、弱く引いて相対的な動きが見えなかっただけでは、固く接合されているとは確定できない。摩擦や、カメラの分解能より小さい動きでも、同じ観測になり得る。その結果は、ある条件のもとでの証拠であって、機構の完全な説明ではない。何が未解決のままかを読み取ることも、学習に含まれる。

帰結から、学ぶ価値のあることへ帰結、問いに対する証拠、学ぶ価値の三段。帰結から証拠への間には問いと測定があり、証拠から価値への間には目的と将来の用途がある。01帰結何が起きたか両方の部品が動いた。02証拠何が分かったか競合する説明を区別できるようにする。03学ぶ価値何が可能になるか別の状況でも使える区別を学ぶ。問い + 測定目的 + 将来の用途
帰結から、学ぶ価値のあることへ帰結、問いに対する証拠、学ぶ価値の三段。帰結から証拠への間には問いと測定があり、証拠から価値への間には目的と将来の用途がある。01帰結何が起きたか両方の部品が動いた。問い + 測定02証拠何が分かったか競合する説明を区別できるようにする。目的 + 将来の用途03学ぶ価値何が可能になるか別の状況でも使える区別を学ぶ。
図 1.帰結から、学ぶ価値のあることへ。測定が問いの証拠になるのは、ある条件のもとでのことであり、問いへの決着を意味するとは限らない。さらに、学ぶ価値は目的と将来の用途にも依存する。矢印は自動的な含意ではなく、学習器が担う仕事を表す。

視点を選ぶ、行動を記録する、変化が分かるフレームを残す。データセットの作り手が、そうした仕事をすでに担っている場合もある。同じ相互作用を繰り返さなくても、その記録から学ぶことはできる。観測を記録したデータセットは、世界の外にあるものではない。世界について得られた観測を、選び取って保存している。

行動できることが加えるのは、次に取得する観測を変えられる能力だ。

新しい技能は、新しい知り方になる

ある問いについて証拠を得られるようになると、次の問いが生まれる。別のことを調べる助けとして、その経験から何が残るのだろうか。私は視覚知識を、それが何を支えるかから考えたい。何が変わらず、何が変わり、条件を変えたときに何が起こるかについての予測である。[2]

その知識は、言葉で表されている必要も、明示的な命題の集まりとして保存されている必要もない。広く使える視覚表現から、事前学習の後にタスクに応じた読み出しを学ぶことで、新しい問いに答えられる場合もある。ここで考えたいのは、手元の観測や、その使い方だけでは足りないときに何ができるかだ。

二つの部品を扱うロボットに戻ろう。片方を安定して保持できるようになることは、操作技能の向上に見える。だがそれは、以前は曖昧だった動きを、情報のある観測へ変えることでもある。物体を回転させる技能は、元のカメラには見えなかった面を明らかにする。鏡の使い方を学べば、新しい視線の経路が得られる。

能力が変えるのは、エージェントが達成できることだけではない。世界から何を学べるかも変わる。

ただし、新しい視点によって現在の不確実性が解消しても、新しい状況での学習能力まで向上するとは限らない。発達につながるには、発見した関係、よりよい測り方、違いが現れる状況のつくり方など、再利用できるものが残る必要がある。運動技能が開くのは、観測へのアクセスだ。そこから得た証拠を学習に使うことは、さらに必要な能力である。

私がつくりたいのは、この発達的な循環だ。知識が行動を可能にし、行動が情報のある観測へのアクセスを開き得る。その観測を解釈することで、さらに知識が得られる。重要なのは、残ったものが元の課題の反復にとどまらず、別の状況で証拠を得たり、読み取ったりする助けになることだ。

技能が、証拠への新しい経路を開く何を知っているか、何ができるか、どんな証拠に届くか、次に何を学べるかを結ぶ循環。固定された循環の反復ではなく、証拠を得る方法自体が発達する。知っていること安定や変化を予測するできること保持する・回す・調べる届く証拠以前は見えなかった関係にアクセスする次に学べること別の問いにも知識を使えるようになる発達する学習の循環学習が、次に学べることを変える。
技能が、証拠への新しい経路を開く何を知っているか、何ができるか、どんな証拠に届くか、次に何を学べるかを結ぶ循環。固定された循環の反復ではなく、証拠を得る方法自体が発達する。知っていること安定や変化を予測するできること保持する・回す・調べる届く証拠以前は見えなかった関係にアクセスする次に学べること別の問いにも知識を使えるようになる学習が、次に学べることを変える。
図 2.技能が、証拠への新しい経路を開く。技能は、情報のある観測へのアクセスを開き得る。その観測を読み取った後に、再利用できる知識や調べ方が残ってこそ、循環が発達につながる。これは発達の可能性であり、無限の成長を保証する法則ではない。

目標は、すべての予測を無差別に疑うことではない。どの予測が誤っているか、あるいは手元の証拠ではまだ定まらないかを調べる方法を学ぶことだ。モデルが最初は高い確信を持っていた予測も、その対象に含まれる。

補足一つの表現から、多くの問いへ

広く再利用できるエンコーダと、能動的に証拠を探す系は補完的である。新しい問いに必要なものが、すでに表現に含まれている場合もある。そのとき必要なのは、別の実験ではなく、別の読み出しかもしれない。

Steerable Visual Representationsは、与えられたテキストで視覚エンコーダを条件づけ、同じ画像の中で異なる概念へ計算を向ける具体例である。[3]この条件づけが変えるのは、手元の情報の使い方だ。プロンプトが事前知識や課題の説明を与えることはあっても、計算の向け先を変えること自体は、写っている場面の新しい測定ではない。問いを変えること、証拠を得ること、知識を更新することは、異なる操作である。

学習が、次に学べることを変える。

確かめられることのすべてに、学ぶ価値があるわけではない

問いを選ぶ仕事を学習器へ渡すと、選び方そのものを誤る可能性が生まれる。簡単な問いをつくり、重要でない規則性を予測し、すでに知っていることを繰り返し確認することばかり、上手になるかもしれない。自分で生成した課題で高い正解率を得ても、その課題の外で役立つことを学んだとは限らない。

これは視覚に限った問題ではない。AIを用いた数学についても、形式的な検証と、その形式化がもとの主張を正しく捉えているかは区別される。また、正しい結果であることと、知的な価値があることも違う。[4]この区別を視覚学習へ引き取るなら、二つの要求がある。証拠が実際に問うていることに関係していること。そして、その問いに答えることが、学習器の労力に見合うことだ。

ある区別を学ぶ価値を考えるには、まず、その区別が何を意味するかを明らかにする必要がある。二つの部品を「固く接合されている」と呼ぶなら、力を加えたときの相対位置について、予測が変わるはずだ。別のラベルを付けるだけでは、その内容を与えたことにならない。

ここで、プラグマティズムが議論に入る。チャールズ・S・パースは『How to Make Our Ideas Clear』で、概念を明瞭にするために、その概念が持ち得る実際的な帰結を考えることを提案した。[5]私はこれを、表現を、確かめ得る予測へ結びつけるための見方として使いたい。その区別が正しいなら、どの条件で、何が違ってくるのか。

これは意味についての考え方であって、どの問いに注意を向けるべきかまで決めるものではない。帰結が明確な区別でも、学習器の目的には関係しないことがある。何が重要かは、その系の目標、感覚と行動の能力、資源にも依存する。Umwelt Representation Hypothesisが強調するのも、この関係である。[6]ただし、便利さが真偽の判定になるわけではない。役に立ちそうな予測でも、間違うことはある。

二つの部品の例に戻ろう。接合状態とは独立に、点灯し続ける表示灯が映っているとする。同じ相互作用の記録から、表示灯についての予測も、部品間の相対運動についての予測も確かめられる。この例で接合状態の区別に役立つのは後者だ。検証できることだけでは、どの標的を学ぶ価値があるかは決まらない。

検証可能な予測でも、学習上の価値は異なる同じ相互作用の記録に対する二つの予測標的を比較する。接合関係を調べるという目的では、接合と独立な表示灯の点灯を予測しても接合の情報は得られない。一方、相対運動の予測は接合関係についての情報になり得る。どちらも検証でき、別の組み立て物への転移は別途確かめる必要がある。目的:部品の接合関係を学ぶ同じ相互作用の記録から、異なる予測標的をつくる。標的 A表示灯の状態t0t1t2記録から検証できるか?はい。観測した状態と比較できる。接合関係について何が分かるか?この例では、接合と独立に点灯する。点灯の予測では、接合を区別できない。標的 B部品間の相対運動ABt0ABt1記録から検証できるか?はい。観測した状態と比較できる。接合関係について何が分かるか?この条件で相対的に動くかを確かめる。一部の説明を除外する証拠になり得る。検証できることだけでは、学ぶ価値は決まらない。判断の改善や、次の学習への再利用は、別に確かめる。
検証可能な予測でも、学習上の価値は異なる同じ相互作用の記録に対する二つの予測標的を比較する。接合関係を調べるという目的では、接合と独立な表示灯の点灯を予測しても接合の情報は得られない。一方、相対運動の予測は接合関係についての情報になり得る。どちらも検証でき、別の組み立て物への転移は別途確かめる必要がある。目的:部品の接合関係を学ぶ同じ記録から、二つの予測標的へ。標的 A表示灯の状態t0t1t2記録から検証できるか?はい。観測した状態と比較できる。接合関係について何が分かるか?この例では、接合と独立に点灯する。点灯の予測では、接合を区別できない。標的 B部品間の相対運動ABt0ABt1記録から検証できるか?はい。観測した状態と比較できる。接合関係について何が分かるか?この条件で相対的に動くかを確かめる。一部の説明を除外する証拠になり得る。検証できることだけでは、学ぶ価値は決まらない。判断の改善や、次の学習への再利用は、別に確かめる。
図 3.検証可能な予測でも、学習上の価値は異なる。モデルの測定結果ではなく、考え方を示す例。表示灯は接合状態と独立だと仮定する。同じ相互作用でどちらの標的も検証できるが、指定した目的への関係は異なる。別の課題では色が重要なこともある。未知の組み立て物への再利用は、ここで実証した結果ではなく、確かめたい可能性である。

ここから私が進めたいのは、知識に価値を与える帰結の中に、将来の学習を含めることだ。片方の部品を保持する方法を知れば、接合関係を観測できるようになる。その関係の理解は、見慣れない機構を調べる助けになるかもしれない。目の前の答えは、得られるものの一部にすぎない。調べ方が別の状況へ転移するなら、その価値は、最初に役立った課題の先まで広がり得る。

こう考えると、発達的な循環の目的を、今日の報酬だけに狭めずに済む。あらゆる違いを集めることが目標ではない。目的と制約のもとで、判断をよくしたり、調べる価値のあることへの道を開いたりする知識を獲得したい。

知識の価値の一部は、それによって次に何を学べるようになるかにある。

探索と学習のための世界モデル

ここで学習器は、どの行動や視点から有用な証拠が得られそうかを選ぶ必要がある。その選び方を経験から直接学ぶことも、予測を使って行動の候補を比較することもできる。後者で、世界モデルの役割が具体的になる。ここでいう世界モデルとは、候補となる行動も含めて、環境とその観測がどう変わり得るかを予測する、学習されたモデルである。

行動する前に考え得る帰結を検討することは、行動選択の助けになる。[7]ここでは、その選択に、学ぶことを直接の目的とする行動も含める。モデルがすでに答えを知っている必要はない。答えの候補によって予測される結果が異なる観測を、見つける助けになればよい。

二つの部品なら、固い接合によって一緒に動く場合と、接触だけで一緒に動く場合について、異なる動き方を予測するかもしれない。その比較から、片方を保持し、相対的な動きを見るという行動が候補になる。ただし、得た結果は、先ほど述べた測定の限界のもとで読み取る必要がある。実際の機構が想像した候補に含まれていなかったなら、観測によって候補自体も修正できなければならない。

同じモデルは、学習用のシミュレーション経験を生成することにも使える。実環境での次の行動を選ぶだけでなく、制御可能な世界モデルで状況のバリエーションを生成し、その中でエージェントが行動を学習・評価する。世界モデルを学習・評価のための資源として使うことへの期待は、ここにある。[8]知識が、さらに知識を獲得するための基盤の一部になる。

ただし、シミュレーション経験からの学習と、現在の実環境から新しい証拠を得ることは、役割が異なる。シミュレーションは、モデルの仮定から何が導かれるかを明らかにし、その仮定のもとで行動を学ぶ助けになる。同じモデルからサンプルを増やしても、モデルの外の場面を正確に捉えているとは確かめられない。シミュレーション内での改善が、実環境へ転移するかを確かめる必要がある。

モデルの予測で行動を選び、観測からモデルを更新する世界モデルが帰結の候補とシミュレーション経験を生成する。提案した行動で実環境と相互作用し、得た観測から予測を更新する。モデル内で生成した帰結と、外部から取得した観測の間に境界がある。モデルの内側行動の帰結を予測する比較する・生成する・学習する実環境との相互作用帰結を観測するOBS測定する・解釈する・更新する問い・行動を選ぶ証拠に照らす想像された帰結 ≠ 世界についての新しい測定
モデルの予測で行動を選び、観測からモデルを更新する世界モデルが帰結の候補とシミュレーション経験を生成する。提案した行動で実環境と相互作用し、得た観測から予測を更新する。モデル内で生成した帰結と、外部から取得した観測の間に境界がある。モデルの内側行動の帰結を予測する比較・生成・学習実環境との相互作用帰結を観測するOBS測定・解釈・更新問い・行動を選ぶ証拠に照らし更新する想像された帰結は、世界についての新しい測定ではない。
図 4.モデルの予測で行動を選び、観測からモデルを更新する。世界モデルは、行動候補の比較や学習用のシミュレーション経験の生成に使える。新しい観測で予測を確かめるには、関係する違いを測れている必要がある。モデルが生成した経験は、実環境の独立した測定ではない。
補足もっと考えることと、新しく見ること

迷路全体が見えているなら、追加の計算によって、証拠からすでに定まっている経路が見つかるかもしれない。一方、扉の状態が隠れていて、手元の手掛かりでは決まらないなら、詳しい未来を描いても、実際の状態が明らかになるわけではない。事前知識が一方の可能性を支持しても、その扉の新しい観測を得たことにはならない。

それでも、生成した候補は、次に何を観測するとよいかを示し得る。基準は、その候補が鮮明かどうかではなく、調べる助けになるかだ。この区別は、特定の生成アーキテクチャには依存しない。

ここで分けているのは、固定された二つのモジュールではなく、二つの時間スケールである。一回の相互作用では、知覚を使って行動を選ぶ。相互作用を重ねる中では、獲得した行動と、その帰結からの学習によって、取得・解釈できる観測の範囲が広がり得る。

一回の判断では、知覚が行動に先立つことがある。学習の時間を通して見れば、行動は知覚に何ができるかを変えられる。

事前学習の先にある汎用性

汎用性を考えるとき、私は二つの能力を求めたい。事前学習で得た知識を新しい課題へ適用することと、その知識では足りないときに新しい経験から適応することだ。既存知識からの汎化と、新しい観測を通じた適応は、異なる能力である。有用な事前学習済み表現は、前者を支え、後者の出発点にもなる。

それでも、二つの状況が手元では同じ観測を与える一方で、答えは異なる場合を考えよう。その観測だけでは、どちらの状況にいるかを特定できない。事前知識が一方の答えをもっともらしくすることも、追加の計算が見落とした含意を引き出すこともある。だが、どちらも、未解決の違いについて新しい観測を得ることではない。これは、その問いに対する証拠の限界であり、受動学習一般に共通する天井ではない。

情報のある観測にアクセスできるなら、次の課題は、新しいタスクのたびに人間がデータ収集と学習標的を設計し直さなくても、必要な観測を取得・解釈できるようにすることだ。目標は、事前学習で得た知識から未知の問いに答えることの先へ進む。その知識では足りないとき、知識の得方自体も適応させる必要がある。

『Visual General Intelligence: A White Paper』の私たちの節では、汎用的な視覚知能(VGI)を、部分的な視覚経験から世界の使える構造を獲得し、新しい課題へ適用する知能として述べた。[9]本稿で広げたいのは、その獲得過程だ。誰かがあらかじめ用意した経験だけでなく、学習器自身が、どんな条件を整えれば必要な観測を得られるかを学ぶことまで含めたい。

補足限界と系譜

これは研究上の立場であり、あらゆる問いが答えられるようになるという保証ではない。センサー、資源、安全上の制約、そして世界そのものが限界を課す。自律的な学習器は、与えられた目的と制約に責任を持ちながら、有用な中間の問いを見つけるべきだ。測りやすいものへ目的をすり替えてよいわけではない。

好奇心駆動の発達学習は、獲得した技能が次の学習への足場になることを、すでに研究している。[10]知覚–行動の視点も、経験の自律的な収集を学習問題の内側に置く。[11]ここで引き受けたいのは、有用な証拠を得て、その意味を読み取ることを、別の状況にも転用できる能力にするという課題である。

この議論は、内部表現を一種類に決めることも、毎回の行動の前に動画を生成することも要求しない。学んだ証拠の得方が、明示的な手順として表されている必要もない。別の状況から学ぶ助けになるかが重要である。

ここに一つ、賭けを置きたい。調べる能力そのものが、別の状況へ転移するはずだ。

ここでいう転移は、新しい物体に対して同じ成功した動作を実行すること以上を意味する。見えるようにする、切り分ける、比べる、試す。その方法を適応させ、不慣れな状況から情報を得られるようになることだ。その調べ方を学んだ時点で、新しい課題が予想されている必要はない。

継続的に物理世界と関わる系にとって、この再利用可能な能力が、汎用性の大きな源になると予想する。広い視覚知識を事前学習で得た系について、さらに強く言えば、未知の物理的な状況への適応では、同種の受動的な観測を増やすだけの場合より、足りない証拠の得方を改善することが重要になると考えている。この優先順位は、さらに確かめるべき主張である。能動的な観測が役立ち得るというだけでは、導かれない。

新しい種類の問いに出会うたびに、人間が新しいデータセットを考案し、標的を指定し、成功をどう読み取るかまで作り直さなければならないなら、学習過程の相当な部分は、なおエージェントの外に残っている。

逆に、新しく学んだ証拠の得方が、元の課題にしか役立たないなら、私が賭けている発達的な循環は、期待よりもずっと弱い。

目標は、何もかも見たことのある機械ではない。知っていることを使って、知らないことを調べ、そこで分かったことを、さらに学ぶ力へ変えられる機械である。

私の賭け

視覚の汎用性を生むのは、何を見てきたかより、何をどう調べられるようになったかである。

参考文献・関連する読み物

  1. 01
    Where Do Good Vision Targets Come From?↩Kohsuke Ide

    学習標的はどこから来るのか。あらかじめ整えられた課題の外で、経験をどう学習につなげるかという出発点。

  2. 02
    Plato Is Not a Space↩Kohsuke Ide

    内部の幾何を先に指定するのではなく、表現が支える予測や応答を通して視覚知識を考える。

  3. 03
    Steerable Visual Representations↩Jona Ruthardt, Manu Gaur, Deva Ramanan, Makarand Tapaswi & Yuki M. Asano · 2026

    与えられた問いや概念に応じて視覚計算を変える具体例。新しい感覚的証拠を得ることとは区別される。

  4. 04
    Mathematical methods and human thought in the age of AI↩Tanya Klowden & Terence Tao · 2026

    §4.4は形式的な正しさと意図した数学的主張への忠実さを、§§4.6–4.7は正しさと広い理解・価値を区別する。本稿の視覚学習への適用は類推であり、同論文で証明された主張ではない。

  5. 05
    How to Make Our Ideas Clear↩Charles S. Peirce · 1878

    考え得る実際的な帰結から概念の意味を明瞭にする方法。知識には将来の学習を可能にする価値がある、という発展は本稿の立場であり、この格率から確立される結論ではない。

  6. 06
    The Umwelt Representation Hypothesis: Rethinking Universality↩Victoria Bosch, Rowan Sommers, Adrien Doerig & Tim C. Kietzmann · 2026

    Box 1と§2。有用な区別や表現の整合を、目標、感覚・行動の能力、資源の制約との関係で考える。

  7. 07
    Coding Is All You Need? Why We Need a World Model!↩Hokin Deng

    行動の前に帰結を検討するという視点。併記した記事は、動画生成中の熟考を論じる。

    Schrödinger’s Cat in Video Models
  8. 08
    World Models at Third Dimension AI↩Christian Rupprecht · 2026

    学習・評価用の制御可能な資源としての世界モデルと、そのモデル自身の学習データを超える汎化の難しさ。本稿が引き取るのはその役割であり、シミュレーション内の経験が現実へ転移する保証ではない。

  9. 09
    Visual General Intelligence: A White Paper↩Hirokatsu Kataoka et al. · 2026

    §2.10。部分的な視覚経験から使える知識を獲得し、予測、想像、再構成、新しい課題へ適用する。

  10. 10
    Intrinsically Motivated Goal Exploration Processes with Automatic Curriculum Learning↩Sébastien Forestier, Rémy Portelas, Yoan Mollard & Pierre-Yves Oudeyer · 2017 / JMLR 2022

    自律的な目標探索と、次の学習への足場となる技能。

  11. 11
    The flavor of the bitter lesson for computer vision↩Vincent Sitzmann · 2026

    知覚–行動ループと、経験を自律的に獲得する問題。