Huh, good point. Appreciate these catches. I am not quite sure what I was thinking here, so for now have deleted the relevant section of the post. (I think there’s some point to be made here about what counts as a fair decision problem, but will need to think about it more.)
I liked that section, I thought it was making an important point. It was useful for my understanding when I first read the post - it distinguishes what you're saying from FDT. FDT definitely thinks *something like* "it’s fair for decision problems to depend on an agent’s entire policy", but doesn't think that it's fair for them to depend on what concepts you're using, like you were saying. (or more properly, it's not even embedded enough to incorporate that any real-life application of it would be using concrete concepts. So it's very confusing to try to apply it to real life "hostile telepaths"-type problems, and your framing was a real improvement for my understanding of those situations).
The fair decision problem thing is a known hard problem, there's a footnote in the FDT paper about how they were unable to find a good precise definition of fairness, and rely on an intuitive notion. But e.g. while Newcomb's Revenge seems intuitively clearly unfair, "hostile telepaths"-type problems seem kind of fair-ish, so your distinction in the section is important.
Maybe add the section back, and add a footnote after "depend on an agent's entire policy" saying "in some kind of non-arbitrary / fair way", and refer to the footnote in the FDT paper (footnote 15), or something like that?
Arguably Troll Bridge is an example of a problem that depends on the agent's reasoning process rather than its policy - but that's in the context of proof-based decision theory.
"Note that this introduces an asymmetry between external actuators that are cooperative vs uncooperative with you. In both cases, they’re reading your mind, and influencing the world based on what they see. But in the cooperative case, this allows you to form stable self-fulfilling beliefs; whereas in the uncooperative case such beliefs are unstable under reflection."
By the Intermediate Value Theorem, a stable belief always exists in any continuous universe. Isn't that in the Logical Inductors paper?
Yes, that seems correct. I should have distinguished not between *some* and *no* stable beliefs, but rather between *many* and *almost none*. Will edit to clarify.
Great post! I suspect that divergence minimisation between preference and belief distributions is a helpful approximate way to formalise this (e.g., https://openreview.net/forum?id=4dc15FtIaD), and I think this is compatible with the belief web + forces framing - it seems fair to say that value and probabilities are reasonably dissociable in the human brain (e.g., https://www.jneurosci.org/content/25/19/4806, https://pmc.ncbi.nlm.nih.gov/articles/PMC7010285/), so I guess one could imagine a dual web system where the drive to reduce inconsistency that naturally operates on beliefs could just be applied to these two webs while modulating the relative precision or "stubbornness" of the preference web to that of the belief web which determines the degree to which inconsistency is resolved by action or belief updating. I'm curious what the arguments would be in favour of getting rid of preference distributions in favour of vector fields.
The basic issue is that I don't know what a preference distribution *is*. Like, suppose you have options A, B, C, and you assign them 30%, 60%, 10% probability respectively. Does this mean that you like B twice as much as you like A? Then it's not really a probability distribution, it's a normalized utility function.
The reason the idea of a preference distribution seems nice is that it would be really elegant for preferences to be the same kind of thing as beliefs, so they could interact via inference. But I think the vector fields thing captures this core intuition without needing the (IMO confused) step of assigning probabilities to goals. And sure, we still need to specify what the drives are, but you can imagine those being very low-level, such that almost all our high-level concepts (including goals) are the same "kind of thing" (namely beliefs).
Thanks for clarifying. Fwiw, the way the preference distribution in active inference is typically constructed from a given (finite) utility function is that instead of just normalising it, you apply a softmax, where the coefficient of the utility in the exponential represents the precision/strength of your preferences. This allows you interpret it as description length minimisation as Wentworth points out (https://www.lesswrong.com/posts/voLHQgNncnjjgAPH7/utility-maximization-description-length-minimization).
I like what I've seen of Leverage's ideas but I haven't seen anything that clearly lays out what connection theory is, so I'm not sure how much that specific framework inspired me.
Nice! :) Yeah, and also the whole framework of the mind as a "belief system", on subsequent pages.
The biggest miss is probably that they didn't connect it to predictive processing / active inference, given the "belief=goals" vibe in 2 and 5. They don't even list it in their enumeration of "previous approaches" on page 3.
Hey Richard — I would like to suggest that nearly everything you want falls straight out of active inference over a sufficiently rich hierarchical generative model with a self-model, and that the two things you present as forcing a new framework aren't doing the work you think:
* Local-but-not-global coherence is just what approximate variational inference is — free energy is the running measure of residual incoherence, and your own footnote 6 about propagation speed is the mechanism.
* Hierarchy is native to ActInf, so all the working around its absence in PDGs/Garrabrant induction is unnecessary.
* Drives and anchors look to me like precisions under other names; goal-credences don't contaminate the posterior once preferences sit in the prior over observations where they belong.
* Actions-as-self-fulfilling-beliefs needs no logical self-reference: the loop is causal and forward in time, so a predictor who models you is just a conditional with a cycle in the graph, which closes numerically as a Brower fixed point of message passing — again, dispensing with the Löb machinery.
The one thing this genuinely gives up is faithful embedding: A's model of B's model of A being the real A. But that requirement is smuggled in rather than argued for, and real agents can't have it, for boring information-theoretic reasons: a finite system cannot carry a lossless copy of itself, let alone of a system carrying a copy of itself. Lossy marginals terminate the regress at depth one and are all any of your motivating cases (Newcomb, sincerity, coups) actually require. So the burden seems to me to be on faithfulness: if the interest is in modelling real-world agents, why demand a property they demonstrably lack?
Happy to spell any of this out properly over email or a call if useful; particularly the precision story, which I think absorbs more of your framework than is immediately obvious.
"It seems like belief webs implicitly implement EDT, which struggles to evaluate hypotheticals without interference from existing beliefs. But might they have emergent FDT/UDT-like properties in a way that gets the best of both worlds? The intuitive link here is that the FDT relies on considering logically impossible hypotheticals without “collapsing” them by propagating the contradiction."
"In active inference/predictive processing, minds are viewed as hierarchical generative models, with each layer of the hierarchy forming new concepts with reference to lower-level concepts."
What is the ontology/ontologies these concepts are formulated in, and how does an ontology get adopted/modified/discarded?
"FDTers think it’s fair for decision problems to depend on an agent’s entire policy (including how it would act in hypothetical scenarios)"
Nope! For example, FDT thinks Newcomb's Revenge is unfair (unless the hypothetical Newcomb's is held in the same ecology).
Huh, good point. Appreciate these catches. I am not quite sure what I was thinking here, so for now have deleted the relevant section of the post. (I think there’s some point to be made here about what counts as a fair decision problem, but will need to think about it more.)
I liked that section, I thought it was making an important point. It was useful for my understanding when I first read the post - it distinguishes what you're saying from FDT. FDT definitely thinks *something like* "it’s fair for decision problems to depend on an agent’s entire policy", but doesn't think that it's fair for them to depend on what concepts you're using, like you were saying. (or more properly, it's not even embedded enough to incorporate that any real-life application of it would be using concrete concepts. So it's very confusing to try to apply it to real life "hostile telepaths"-type problems, and your framing was a real improvement for my understanding of those situations).
The fair decision problem thing is a known hard problem, there's a footnote in the FDT paper about how they were unable to find a good precise definition of fairness, and rely on an intuitive notion. But e.g. while Newcomb's Revenge seems intuitively clearly unfair, "hostile telepaths"-type problems seem kind of fair-ish, so your distinction in the section is important.
Maybe add the section back, and add a footnote after "depend on an agent's entire policy" saying "in some kind of non-arbitrary / fair way", and refer to the footnote in the FDT paper (footnote 15), or something like that?
"Robust Cooperation in the Prisoner's Dilemma: Program Equilibrium via Provability Logic" talks about this kind of problem some.
Arguably Troll Bridge is an example of a problem that depends on the agent's reasoning process rather than its policy - but that's in the context of proof-based decision theory.
"Note that this introduces an asymmetry between external actuators that are cooperative vs uncooperative with you. In both cases, they’re reading your mind, and influencing the world based on what they see. But in the cooperative case, this allows you to form stable self-fulfilling beliefs; whereas in the uncooperative case such beliefs are unstable under reflection."
By the Intermediate Value Theorem, a stable belief always exists in any continuous universe. Isn't that in the Logical Inductors paper?
Yes, that seems correct. I should have distinguished not between *some* and *no* stable beliefs, but rather between *many* and *almost none*. Will edit to clarify.
Great post! I suspect that divergence minimisation between preference and belief distributions is a helpful approximate way to formalise this (e.g., https://openreview.net/forum?id=4dc15FtIaD), and I think this is compatible with the belief web + forces framing - it seems fair to say that value and probabilities are reasonably dissociable in the human brain (e.g., https://www.jneurosci.org/content/25/19/4806, https://pmc.ncbi.nlm.nih.gov/articles/PMC7010285/), so I guess one could imagine a dual web system where the drive to reduce inconsistency that naturally operates on beliefs could just be applied to these two webs while modulating the relative precision or "stubbornness" of the preference web to that of the belief web which determines the degree to which inconsistency is resolved by action or belief updating. I'm curious what the arguments would be in favour of getting rid of preference distributions in favour of vector fields.
The basic issue is that I don't know what a preference distribution *is*. Like, suppose you have options A, B, C, and you assign them 30%, 60%, 10% probability respectively. Does this mean that you like B twice as much as you like A? Then it's not really a probability distribution, it's a normalized utility function.
The reason the idea of a preference distribution seems nice is that it would be really elegant for preferences to be the same kind of thing as beliefs, so they could interact via inference. But I think the vector fields thing captures this core intuition without needing the (IMO confused) step of assigning probabilities to goals. And sure, we still need to specify what the drives are, but you can imagine those being very low-level, such that almost all our high-level concepts (including goals) are the same "kind of thing" (namely beliefs).
Thanks for clarifying. Fwiw, the way the preference distribution in active inference is typically constructed from a given (finite) utility function is that instead of just normalising it, you apply a softmax, where the coefficient of the utility in the exponential represents the precision/strength of your preferences. This allows you interpret it as description length minimisation as Wentworth points out (https://www.lesswrong.com/posts/voLHQgNncnjjgAPH7/utility-maximization-description-length-minimization).
Wow, really interesting, thanks.
Were you also inspired by Leverage's connection theory here? This makes me understand what Geoff might be trying to get at for the first time.
I like what I've seen of Leverage's ideas but I haven't seen anything that clearly lays out what connection theory is, so I'm not sure how much that specific framework inspired me.
Makes sense - yeah, I was being imprecise since they don't even use the term anymore, sorry. I'm going off of this PDF, (https://cdn.prod.website-files.com/66fcf851dc59e7cd95ec409d/67519c51d3bce3f33da46068_Summary%20of%20Leverage%20Introspection%20Research%20-%20Export%20%231.pdf), page 11. But yeah the connection is a bit tenuous.
Oh, no, this is super relevant. I just hadn't seen their ideas explained this way before. For easy reference:
"Our researchers found that the relevant parts of the mind could be described by the following rules:
1. A person’s beliefs update:
• ...in order to explain their sensations,
• ...elegantly, i.e., towards mutual coherence,
• ...locally, i.e., towards local rather than global elegance, and
• ...only where they pay attention.
2. A person always believes their basic goals will be achieved.
3. A person’s basic goals do not change.
4. A person has exactly those concepts included in their beliefs.
5. People act in accordance with what they believe will happen.
6. A person’s intention is included in their attention."
This seems extremely prescient, I guess they just never managed to connect it to the agent foundations stuff (or vice versa).
Nice! :) Yeah, and also the whole framework of the mind as a "belief system", on subsequent pages.
The biggest miss is probably that they didn't connect it to predictive processing / active inference, given the "belief=goals" vibe in 2 and 5. They don't even list it in their enumeration of "previous approaches" on page 3.
Hey Richard — I would like to suggest that nearly everything you want falls straight out of active inference over a sufficiently rich hierarchical generative model with a self-model, and that the two things you present as forcing a new framework aren't doing the work you think:
* Local-but-not-global coherence is just what approximate variational inference is — free energy is the running measure of residual incoherence, and your own footnote 6 about propagation speed is the mechanism.
* Hierarchy is native to ActInf, so all the working around its absence in PDGs/Garrabrant induction is unnecessary.
* Drives and anchors look to me like precisions under other names; goal-credences don't contaminate the posterior once preferences sit in the prior over observations where they belong.
* Actions-as-self-fulfilling-beliefs needs no logical self-reference: the loop is causal and forward in time, so a predictor who models you is just a conditional with a cycle in the graph, which closes numerically as a Brower fixed point of message passing — again, dispensing with the Löb machinery.
The one thing this genuinely gives up is faithful embedding: A's model of B's model of A being the real A. But that requirement is smuggled in rather than argued for, and real agents can't have it, for boring information-theoretic reasons: a finite system cannot carry a lossless copy of itself, let alone of a system carrying a copy of itself. Lossy marginals terminate the regress at depth one and are all any of your motivating cases (Newcomb, sincerity, coups) actually require. So the burden seems to me to be on faithfulness: if the interest is in modelling real-world agents, why demand a property they demonstrably lack?
Happy to spell any of this out properly over email or a call if useful; particularly the precision story, which I think absorbs more of your framework than is immediately obvious.
"a finite system cannot carry a lossless copy of itself, let alone of a system carrying a copy of itself"
Can you elaborate on this point? Aren't quines a counterexample of the former, and quine-relays a counterexample of the latter?
"It seems like belief webs implicitly implement EDT, which struggles to evaluate hypotheticals without interference from existing beliefs. But might they have emergent FDT/UDT-like properties in a way that gets the best of both worlds? The intuitive link here is that the FDT relies on considering logically impossible hypotheticals without “collapsing” them by propagating the contradiction."
I'm still not entirely sure what you mean by this. You're aware that MIRI researchers have leant more towards EDT-style counterfactuals over time (vindicating Wei Dai from way back in the day)? Abram Demski: (https://www.lesswrong.com/posts/yXfka98pZXAmXiyDp/my-current-take-on-counterfactuals), Scott Garrabrant: "EDT is obviously philosophically correct" (https://www.lesswrong.com/posts/FBbHEjkZzdupcjkna/miri-op-exchange-about-decision-theory-1). But maybe I'm misunderstanding you.
Wait, yeah I think I am misunderstanding you. But I'm still confused, e.g. what are the two things in "both worlds"?
Do I recognize the smell of Anima Labs in this essay?
Great post - feels like it point to something that explains multi-agent systems in a scale-free manner.
"In active inference/predictive processing, minds are viewed as hierarchical generative models, with each layer of the hierarchy forming new concepts with reference to lower-level concepts."
What is the ontology/ontologies these concepts are formulated in, and how does an ontology get adopted/modified/discarded?