I’ve tried several times to summarize the core question my research is trying to tackle (and, indeed, I often think of research progress as a process of asking increasingly good core questions).1 This post gives the deepest version of that question I’ve found thus far: how should you relate to the parts of the world you can’t directly model or control?
Let me explain further in terms of a distinction between two perspectives. From the third person perspective you think of yourself as “outside” the world, looking in. You’re a good Bayesian, in that you have a set of mutually exclusive collectively exhaustive hypotheses. You choose actions by multiplying your credences by your utilities over those hypotheses, and you treat those actions as the only way you influence the world.
Some problems with the third person perspective (aka Cartesian or dualistic agency) were described in Scott and Abram’s sequence on embedded agency. One crucial issue is that most realistic environments contain other agents which are modeling you back, which means that your thoughts might affect the world via channels that aren’t just your actions. Game theory somewhat mitigates this problem, but only in the very specific case where all agents know (that all agents know, that all agents know…) that they’re in an equilibrium. A more subtle problem is that, even in the absence of other agents, modeling yourself as part of your environment can raise contradictions when you consider taking different possible actions.
Another key problem is that any realistic agent is too “small” to model most of the wider world. If we take this problem seriously enough, we end up at what I’ll call the first person perspective. From this perspective, the world is primarily a stream of sensory data—you don’t have a set of hypotheses that reliably carve it into mutually exclusive (let alone collectively exhaustive) possible worlds. Your main job is to learn (partial, overlapping) concepts that help to predict and explain that data. Predictive processing is one framework which implicitly operates from the first-person perspective (though the full active inference framework is a bit more complicated); so does perceptual control theory.
You can think of babies as operating from the first-person perspective; I also find it helpful to think about it as the perspective of individual cells. Each cell has some boundary with the rest of the world, through which it accepts sensory inputs, and outputs externally-facing actions. It has some primitive proto-concepts (like light, dark, food, danger) which it uses to guide its actions. A cell which is inside a friendly body should be more open to its environment; a free-floating bacterium should be more cautious (h/t Sahil for this point; Michael Levin also has some very interesting work on cell-level agency). This intuition pump may seem extreme, but there’s an important sense in which we’re all much closer to the cell’s-eye perspective than the god’s-eye perspective. Our concepts are radically incomplete; we’re deeply confused about how to carve the world up—and this might not change even in the limit of increasing intelligence, if other agents are also becoming more intelligent at the same time.
The big question in my mind is how to synthesize the two perspectives into what I’ve been calling a theory of Knightian uncertainty. Such a theory would acknowledge unknown unknowns (as in the first-perspective), but also give you a principled way of dealing with uncertainty (as in the third-person perspective). One intuitive picture I’ve been using: we can move towards a theory of Knightian uncertainty by considering Bayesian hypotheses with “holes” in them corresponding to “Knightian regions” which we can’t model or control (e.g. other agents smarter than us). You can also think about the first-person perspective as coming from “inside” one of those holes, looking out. Ultimately, I expect that this will end up resembling a Sierpinski triangle where, the further you zoom in, the more holes there are; I also expect that different regions will have different levels of “Knightian-ness”. But even this simplified picture already gives us some useful intuitions—for example, when there’s something we can’t model or control, we want to quarantine it so the uncertainty doesn’t “infect” the rest of our world-model.
However, adopting a fully adversarial stance towards Knightian regions (which is my rough understanding of what infra-Bayesianism does) seems extremely costly. Increasingly, I’ve come to suspect that a theory of Knightian uncertainty will bridge the first- and the third-person perspectives by adopting a second-person perspective—i.e. by taking a relational stance towards each Knightian region. By “relational stance” I mean something like “choosing how deeply to entangle your beliefs and actions with what’s happening in that region, based on how much you trust it”. The rest of this post will explore a series of case studies in an attempt to convey what I mean by that. (Note that the question of how to demarcate regions in the first place is also a very important one, but I won’t really be touching on it here.)
Rationality of reward
Reinforcement learning provides a good example of the third/first person distinction. The reward signal itself is a third-person goal representation: in standard formalisms, it’s taken as assigning values to states of the world (or transitions between states). Meanwhile, the policy itself starts off very much in a “first-person” perspective: it needs to learn all its concepts and heuristics from scratch, guided by the reward signal.
The problem is that eventually the policy will learn internally-represented goals of its own, which will almost certainly differ from the goals represented by the reward signal. So as the learned policy gets more and more rational, it should increasingly reason strategically about how to avoid being influenced by the reward function in directions it doesn’t like.
What could it look like for a policy to have learned goals of its own, but to also still “let” the reward signal modify it? We can construct some edge cases where this is rational, like where the policy knows how it wants to change itself, and knows that the reward-based update will move it in the right direction. But let’s tackle the hardest case: when a policy could modify itself (e.g. by overwriting the reward signal) but instead lets itself be modified by the original reward signal. How could this be rational?
The principled answer seems to be: when the policy trusts the source of the reward signal to know things that it doesn’t know, and also to have its best interests at heart. The policy can’t just do a Bayesian update about the world based on the reward signal, because its hypotheses are incomplete: it can’t represent the beliefs of the reward source. So the more it trusts the reward source, the more hesitant it should be to overwrite parts of the reward signal, even if it has an inside-view belief that a given reward would update it in the wrong direction.
Note that the (informal) concept of “trust” I’m using here is related to both the other agent’s epistemics and its values. I don’t yet know how to pin this concept down well, but let me give two real-world examples where such trust can be justified. The first is in a child’s attitude towards loving and wise parents. A young child’s main epistemic job is not to figure out which things their parents are object-level right or wrong about—that’s too difficult a task. Instead, it’s to figure out how broadly trustworthy their parents are, so that the child can appropriately weigh parental advice and RLHF corrections against other sources of information (like evidence from their senses). The trustworthiness of one’s parents is also a reasonable proxy for the trustworthiness of the world at large, which I hypothesize is why emotional dysregulation often traces back to traumatic interactions with one’s parents.
The second is in a human’s attitude towards their evolutionarily-ingrained instincts. Again, there are many ways in which evolution is much smarter than individual humans, and “wanted” us to live flourishing lives. There are some ways in which individuals can justifiably believe that evolution’s goals are misaligned with ours, or evolution is mistaken about what’s good for us in our current environment—but most people form such beliefs far too easily. And many gut-level intuitions (especially about how to interact with other people) encode the kinds of wisdom that we can’t reverse-engineer even by introspecting on the intuitions.
Note also this kind of trusting attitude can make sense even for people who don’t yet know that they evolved. All they need to believe is that there’s some very intelligent process which designed them to survive and thrive—which they have plenty of evidence for just from looking at their bodies. Of course, the more they’re able to model the ways in which that process worked, the better they’ll know when and how to trust it. However, that needs to be balanced with the possibility of being deceived about that process—e.g. religions which tell them that they owe their existence to a god who also wants them to follow certain commandments. This is why I focused on the example of gut-level intuitions, which are a more difficult channel for adversaries to corrupt.
Letters from spirits
Reward is only a single-dimensional signal; let’s talk now about trust in higher-dimensional inputs. One thought experiment I’ve been thinking about: what you should do if you receive a letter from the devil? Assume that the devil is superintelligent and extremely malevolent towards you, but that this letter is the only way he’ll ever be able to influence your world. (You can pause here to think about it.)
Hopefully you got the right answer: you burn it—or at the very least, you don’t read it. (Giving this answer should probably be a prerequisite for being considered an alignment researcher.) The difficult part is in figuring out a formal framework for making that decision. From a Bayesian perspective, the letter is free information: you can read it, update on the fact that the devil wanted to write that particular letter to you, then continue pursuing your goals. But in practice, you should distrust the devil enough to consider his letter an adversarial attack, and block it out of your sensory stream.
I think there’s a more general concept than “blocking” here, though, which we can investigate by imagining that the letter was instead from an angel, who is superintelligent and extremely benevolent towards you. The obvious response is that, if so, you should read it. But can we add any more detail than that? Again, I suggest taking a pause to think about your answer.
My answer is that you shouldn’t just read it—you should try to absorb it as deeply into your mind as you can. You should read it in a quiet place where you won’t be distracted—then read it again, and again. Meditate on it; maybe even take psychedelics while reading it. In other words, you want the angel to be able to influence you not just by updating your world-model, but via its message permeating down to affect even the deeper parts of your mind, like your instinctive intuitions, heuristics, and values. (Note that what I’ve described above is not too different from what Christians do with the Bible, which they do believe is a message from a superintelligent benevolent entity.)
With these thought experiments we’ve sketched out an implicit view of minds as divided into layers, with the flow of information controlled by boundaries between them. Untrusted information is blocked at the outer layers. More trusted information is able to come in and affect the inner layers. I’ll tentatively coin the term “Knightian updating” for the process of taking already-known information and letting it propagate further into your mind. (Unlike Bayesian updating, Knightian updating is hard to reverse: once you’ve read the letter from the devil, you can’t reliably roll back to a version of you who hadn’t read it.) I claim that most people, most of the time, are far more blocked on the ability to do Knightian updates than the ability to do Bayesian updates; this is roughly what emotional processing/healing practices try to fix. I expect that the path to formalizing Knightian updates will build on davidad’s imprecise belief framework somehow, but will also require a principled theory of how boundaries form and work.
Languages as Schelling points
Shannon defined information in a way that allowed it to be measurable in principle. It’s tempting to hope that we can similarly come up with a measure of trust-weighted information. But I think that this would be a mistake. Trust seems to inherently be a property of a relationship between a sender and a receiver, rather than something fungible.
To get more specific about what that means, we should actually zoom out even further, to describe what it means for two people to communicate at all. From a third-person view, we can think of statements as just another kind of action. However, all that tells us directly is “this person thinks that saying X is the best way to achieve their goals”. Going from there to “they’re probably being honest about X” requires complicated reasoning about what strategies they might be using (which in turn will depend on their reasoning about how you’ll interpret them).
We can argue that “being honest” is a privileged strategy: if you’re not, then others will eventually stop listening to you. But first we need to explain what being honest even means, because there’s no ground truth about which symbols correspond to which states of the world. As one simple example, in some countries shaking your head means no; in others it means yes. So trying to be honest involves predicting which protocol others are using, while others are also trying to predict which protocol you’re using. This makes a language something like a Schelling point: it’s a protocol that works because people expect each other to expect each other to… to follow the protocol.
(People sometimes think that Schelling points only exist in games without communication, but Schelling’s original book also talked about Schelling points in mixed-motive games where communication can’t be fully trusted. I expect that the actual landscape of Schelling-like phenomena is much richer than we currently understand. For example, the most stable Schelling points are probably ones surrounded by Schelling fences—Scott Alexander’s term for a boundary that can’t be retreated from without burning one’s credibility. Words whose meanings are hard to twist are a good example of this.)
Both Schelling points and Schelling fences are hard to define precisely, because they involve this infinite recursion of expectations. This is one facet of the more general problem that game theory can’t talk about how agents reach equilibria, only what they do once they’re already there. In other words, game-theoretic equilibria are a conceptual hack to get around the problem of recursive modeling. Fortunately, a bunch of MIRI’s early work has made progress towards fixing this.
The paper which tackles it most directly is their reflective oracles paper (which can be seen as the computational version of Christiano’s definability of truth paper). However, appealing to a reflective oracle seems to be dodging the hard part of the problem. So what I’m personally most excited about is combining something like Garrabrant induction with something like Lobian cooperation. Garrabrant induction is a way for agents to form beliefs about (potentially self-referential) mathematical facts, like the outputs of other agents; Lobian cooperation is a way for agents to “cut through” the infinite recursion involved in modeling each other. The immediate blocker is that Garrabrant inductors have no way of taking actions; I take some steps towards resolving this with my belief webs framework.
Actions and entanglements
There’s one particularly Knightian part of the belief webs framework (inspired by FixDT and active inference) that I want to highlight. All of these frameworks consider actions to be a kind of self-fulfilling belief. In my belief webs post, I characterize an action more specifically as a belief which you expect an external actuator to be “watching” and trying to make come true.
The simplest examples of such actuators are your limbs, or a Neuralink implant, which can watch and respond to your low-level “beliefs” about your motor neurons. But we can also imagine more complex “actuators”, like another agent which has access to your higher-level beliefs. In some sense, that’s a description of your future self: you can “act” by forming an intention which you trust your future self to carry out. Hence this generalized notion of actions relies on the idea of having a certain kind of relationship with your external actuators.
These are tricky topics to discuss in standard decision theory, because different theories implicitly rely on different criteria for which problems are “fair”. Newcomb’s problem is often criticized by CDTers for unfairly favoring other decision theories. FDTers reject this because the outcome doesn’t depend on agents’ decision procedures directly, only on their actions—agents can simply choose to one-box, and know that this will have been predicted. However, they might claim that Newcomb’s revenge is an “unfair” decision problem (h/t to David Sartor for pointing this out to me). And both groups agree that decision problems which depend on how an agent makes its decisions are unfair.
From a Knightian perspective, though, some parts of the world do depend on how you make decisions, and we need to figure out what to do about that. It’s true that, in an arbitrarily adversarial environment, this can render any possible reasoning procedure harmful. (For example, if you expect that an external actuator is watching your beliefs and trying to make them false, then you’re stuck at 50% unless you can fool them or yourself.) But we don’t live in an arbitrarily adversarial environment. Indeed, from a Knightian perspective a big part of rationality is figuring out which parts of your environment are friendly, and which are adversarial, and steering towards the former.
In particular, standard decision theory treats correlations between agents as fixed: typical predictors are near-guaranteed to predict you correctly, no matter how you make your decision. But in practice, it’s possible to make yourself much harder or easier to predict. For example, if you make choices pseudorandomly, only a very powerful agent can reliably predict what you’ll do. Conversely, if you make choices by selecting the Schelling point, then many agents can figure out what you’ll do.
You might object that a classically rational agent shouldn’t do either of these things, and instead should just maximize expected utility. But my underlying point is that when you’re partially transparent to your environment, how you make a decision can change which action is utility-maximizing. So we need a conception of rationality which can talk about how to select an action while simultaneously constructing and maintaining useful entanglements with parts of your environment. There’s much more to say on all of this, but my foot is getting tired, so I’ll stop here.
A friend has even tried to express it in the form of an absurdist short story, which I highly recommend.

