Lately I’ve been writing a lot about the sociopolitical considerations that economic-style reasoning misses. I view this work as continuous with my agent foundations research; in this post, I lay out some of the connections I see between sociology and agent foundations. I’ll particularly focus on attempts to pin down theories of rational social and political behavior, as an extension to (or replacement of) the theory of rational economic behavior originally laid out by Von Neumann and Morgenstern.
This post will of course be far from exhaustive. However, it’s worth noting that attempts to develop big-picture frameworks for social theory have grown surprisingly sparse. Berger describes the field as being deformed from two directions since the mid-1900s: on one side, by the “methodological fetishism” of “sociologists using increasingly sophisticated methods to study increasingly trivial topics”; and on the other side, by the “cultural revolution [that] sought to transform sociology from a science into an instrument of ideological advocacy”. Hence many of the thinkers I’ll be discussing were officially based in economics or political science departments.
This post has three sections, each focused on a different topic: social and game-theoretic equilibria; the internalization of norms and values; and collective behavior. In a follow-up post, I’ll talk about more mainstream sociological preoccupations of capitalism and class, and how they relate to rationalist thinking.
Social equilibria vs game-theoretic equilibria
One important concept is that of a “social equilibrium”: a social system which, when disturbed, would tend to return to its original state. This concept was pioneered by Vilfredo Pareto in the early 1900s, and further developed by Talcott Parsons in the mid-1900s. While I haven’t investigated Parsons deeply, his work seems important—Herbert Gintis claimed (in 2009) that “sociological theory has atrophied since the death of Talcott Parsons in 1979”.1
However, progress towards a formal definition of social equilibria has been slow, because it’s hard to model disturbances to a game-theoretic equilibrium. The concepts of trembling hand equilibria, evolutionarily stable strategies and stochastically stable equilibria are three attempts to describe more robust versions of Nash equilibria. However, they only model arbitrary or random deviations from equilibria, rather than ones deliberately caused by players themselves. This leads to problems when talking about counterfactual behavior, as Kenneth Binmore articulated in 1997:
What keeps a rational player on the equilibrium path is his evaluation of what would happen if he were to deviate. But, if he were to deviate, he would behave irrationally. Other players would then be foolish if they were not to take this evidence of irrationality into account in planning their responses to the deviation. A formal model that neglects what would happen if a rational player were to deviate from rational play must therefore be missing something important, no matter how elaborately it is analyzed. However, Aumann, for example, is insistent that his conclusions say nothing whatever about what players would do if vertices of the game tree off the backward-induction path were to be reached. But, if nothing can be said about what would happen off the backward-induction path, then it seems obvious that nothing can be said about the rationality of remaining on the backward-induction path.
In his 2009 book The Bounds of Reason, Herbert Gintis claims that the problem comes from the common knowledge of rationality criterion, arguing that backward induction gives counterintuitive results because common knowledge of rationality is too strong an assumption.2 There is a literature on forward induction which aims to address the same problem Binmore describes, but which seem to rely on similarly strong criteria, such as rationality and common strong belief of rationality.
These assumptions, as well as the concept of Nash equilibria in general, seem to me to be hacks to get around the core sociological problem that Parsons called “double contingency”: the idea that communication (and strategic interaction more generally) requires people to model another person who is simultaneously modeling them back. Intuitively speaking, double contingency is the reason why we can’t talk about how rational agents converge to a game-theoretic equilibrium (or which one they converge to), only what happens when they’re already at an equilibrium. Two other prominent explorations of this same core idea are Keynesian beauty contests (1936) and Schelling points (1960) (see also Scott Alexander here and here, and my discussion of languages as Schelling points).
Harsanyi identified a very similar issue in a 1967 paper:
It seems to me that the basic reason why the theory of games with incomplete information has made so little progress so far lies in the fact that these games give rise, or at least appear to give rise, to an infinite regress in reciprocal expectations on the part of the players.
Harsanyi’s response was to develop a notion of types which defined players’ beliefs about the whole infinite regress. Meanwhile Aumann’s approach (in 1976) was to formalize the notion of common knowledge, which allowed him to prove his agreement theorem. And Gintis argued that we should interpret social norms not as Nash equilibria but rather as correlated equilibria (in which a “choreographer” sends a signal before the game which agents can use to coordinate). I haven’t investigated any of these approaches deeply enough to confidently evaluate them, but none pattern-match to the kind of insight which might make us much less confused about double contingency. Even Leike et al.’s solution to the grain of truth problem required reflective oracles that could be accessed by all players.
Instead, I’m more excited about research on Lobian cooperation, which provides an interesting and counterintuitive model of agents cutting through the infinite recursion of modeling each other. It was first discovered by Slepnev, written up by Barasz et al., and extended by Critch. In its original form, it only applied to proof-based agents who are given each other’s source code. However, we’d ideally like a version of it applicable to agents with probabilistic beliefs about the world and each other (like Garrabrant inductors). Some work that moves in this direction has been done by Payor, Critch, and Demski.
My longer-term hope is that work in this direction helps build up a formal understanding of the kind of intersubjective social epistemology that Gintis calls for in the quote below:
The most fundamental failure of game theory is its lack of a theory of when and how rational agents share mental constructs. The Bayesian rational actors favored by contemporary game theory live in a universe of subjectivity and instead of constructing a truly social epistemology, game theorists have developed a variety of subterfuges that make it appear that rational agents may enjoy a commonality of belief (common priors, common knowledge), but all are failures.
Internalizing norms
Parsons took a different approach to the problem of double contingency, by focusing on agents which internalize norms and values. He discusses this process in his 1951 book The Social System, and argues that it’s necessary to hold together social equilibria:
It is only by virtue of internalization or institutionalized values that a genuine motivational integration of behavior in the social structure takes place, that the “deeper” layers of motivation become harnessed to the fulfillment of role-expectations. It is only when this has taken place to a high degree that it is possible to say that a social system is highly integrated, and that the interests of the collectivity and the private interests of its constituent members can be said to approach coincidence.
This integration of a set of common value patterns with the internalized need-disposition structure of the constituent personalities is the core phenomenon of the dynamics of social systems. That the stability of any social system except the most evanescent interaction process is dependent on a degree of such integration may be said to be the fundamental dynamic theorem of sociology. It is the major point of reference for all analysis which may claim to be a dynamic analysis of social process.
It is the significance of institutional integration in this sense which lies at the basis of the place of specifically sociological theory in the sciences of action and the reasons why economic theory and other versions of the conceptual schemes which give predominance to rational instrumental goal-orientation cannot provide an adequate model for the dynamic analysis of the social system in general terms. It has been repeatedly shown that reduction of motivational dynamics to rational instrumental terms leads straight to the Hobbesian thesis, which is a reductio ad absurdum of the concept of a social system. This reductio was carried out in classic form by Durkheim in his Division of Labor. But Durkheim’s excellent functional analysis has since been enormously reinforced by the implications of modern psychological knowledge with reference to the conditions of socialization and the bases of psychological security and the stability of personality, as well as much further empirical and theoretical analysis of social systems as such.
I think that Parsons concedes too much here. Even if internalization of norms is an important driver of social behavior, it’s still possible in principle for that to coexist with rational goal-orientation. Indeed, much work in AI alignment has been motivated by the idea of agents rationally modifying themselves, including by making commitments over time. On the other hand, value change has always been hard to reason about in economic frameworks, so perhaps it was reasonable for Parsons to interpret norm internalization as being opposed to (the dominant conception of) rationality.
However, my sense is that the econ-minded parts of the field went too far in the other direction over the following decades—as I’ll illustrate with James Coleman’s 1990 book Foundations of Social Theory. It starts with a declaration that he’ll interpret humans as utility maximizers, and follows with 1000 pages of analysis of social phenomena from that lens. In doing so, Coleman showcases the limitations of economic-style thinking about social behavior. For example, his characterization of norms involves individuals holding the right to control decisions about others’ actions. However, in doing so he appeals to implicit property rights over decisions, punting our confusion about norms to the level of how such rights are created and maintained.
There’s been some notable work since Coleman. In particular, Binmore argued that norms should be viewed as equilibrium-selection devices (building on the game theory folk theorem, which established that almost any outcome can be a Nash equilibrium in indefinitely repeated games). Claude also pointed me to Bicchieri’s characterization of norms as conditional preferences and Skyrms’ characterization of signaling conventions. There’s also a bunch of work about norm-following as reputation management; the paper that I’m most familiar with is Paul de Font-Reaulx’s Do expected utility maximizers have commitment issues? I haven’t explored any of these very deeply, but my sense is that all of these approaches either restrict how agents reason about each other or take shared expectations as given.
I’m more interested in research which tries to “get underneath” the standard assumptions of game theory. Hence I’ll return to Coleman—while most of his book builds on existing frameworks, he does question them in Chapter 19, on the construction of the self. Coleman identifies that the standard assumption of a unitary self is suspicious. One alternative he proposes is to distinguish between the “object self” (which has interests and preferences) and the “acting self” (which makes decisions on behalf of the object self). Outside the context of the individual, Coleman describes this as similar to the relationships between children and parents, or shareholders and managers. In the AI context, this is similar to the terminal vs instrumental goal distinction; and also to the relationships between humans and corrigible AIs.
However, I think that progress towards a theory of sociopolitical rationality will require moving beyond this distinction. An approach I like more is to think of agents as coalitions of subagents; Coleman’s term for this is “corporate actors” (by which he means not just corporations but any system that “incorporates” many actors). Ainslie’s picoeconomics takes a similar approach. A major preoccupation of such coalitional agents is resolving internal inconsistencies; Coleman mentions Heider’s “balance theory” of psychology, and Lecky’s self-consistency theory of personality, as both focusing on this. I also characterized identity shifts in terms of people “working to find or maintain [consistent] internal narratives” in my post Higher Education as Class Commitment.
One difficulty in developing these models is that, until recently, we didn’t have many formal frameworks for describing internal inconsistency. Richardson’s probabilistic dependency graphs, Garrabrant’s inductors, and Davidad’s imprecise beliefs, are three recent frameworks for doing so. However, my sense is that they still don’t capture important aspects of internal conflict—e.g. dynamics where individuals try to block out information, which seem important for explaining how egos and identities work. To do so, I expect we’ll need to move beyond the Bayesian expected utility maximization paradigm, towards theories of how to act rationally under Knightian uncertainty.3
Collective behavior
The frameworks for internal inconsistency I mentioned above still focus on the single-agent setting, despite value change typically being induced by multi-agent interactions. My sense is that the best thinking on this topic is mostly non-formal, and weaves together insights from a variety of domains—like Girard on mimesis, Hoffman on depravity, Spence on self-deception, and Lacan on desire.4 And of course we can trace related ideas further back to Freud and Jung. Unfortunately, I’m not aware of many people deeply familiar with continental thought who are writing with the kind of clarity I typically look for (the writers of Pfeilstorch being a notable exception). I’ve personally tried to link some of these ideas to the game theory framework in this talk and this post.
To discuss the standard game-theoretic perspective on collective behavior, I’ll start with the other chapter from Foundations of Social Theory I found most interesting: Chapter 9, on collective behavior. One thing I like is that Coleman’s central examples of game-theoretic interactions are more realistic than most thought experiments. Rather than a Prisoner’s Dilemma he talks about the “game” of leaving a theater after someone yells fire. The best outcome is everyone walking calmly out; the worst outcome is a stampede. This provides a visceral intuition for how cooperation can be a flexible and continuous process (as I discuss in this post).
Coleman understands the intuition that there’s something irrational about the defect-defect equilibrium. His solution (reminiscent of his definition of norms discussed above) is to talk about individuals transferring control of their actions to others. This lets him describe a coordination mechanism similar to FDT-style correlations (although Coleman explicitly denies that his framework is relevant to single-shot games). Though his mathematical framework doesn’t look particularly promising, one thing I like about this approach is that the idea of transferring control to others makes correlated behavior something which agents decide to induce, whereas FDT-style correlations are taken as given. So Coleman’s approach is in some ways closer to my own nascent notion of “entanglement”.
Elsewhere in his book Coleman also discusses an earlier cluster of researchers who he calls the “power theorists”. The ideas I’ve personally found most useful related to this line of thinking are threshold models, information cascades, preference falsification, and coups as coordination games. My sense is that these provide good diagnoses of the problem of herd behavior (and that a better understanding of information cascades would have been very valuable for EA discussions of epistemic modesty) but don’t make much progress towards a theory of group rationality which would prevent this problem.5
The problem of group rationality is related to the problem of credit assignment in multi-agent reinforcement learning—a very difficult problem, since the effects of each agent’s behavior are mediated by other agents. My sense is that MARL researchers, like game theorists, don’t know how to describe the process of rationally choosing between possible equilibria, and instead focus on running agents and seeing what happens in practice. Conceptual clarity about multi-agent learning processes therefore seems rare—though two thinkers whose contributions I appreciate on these topics are Andrew Critch and Jan Kulveit.
It’s worth noting that throughout this post, I’ve been particularly positive about the contributions of agent foundations researchers. This might just be an ingroup bias in my research taste. But my sense is that many agent foundations researchers are searching for a very deep and principled understanding of intelligence, to a much greater extent than researchers in other domains.6 I credit Yudkowsky for anchoring the field of agent foundations towards the pursuit of fundamental knowledge (e.g. in this post). A central example: agent foundations is much more rigorous about self-reference than other frameworks for describing intelligence (like ML theory, or active inference), as I explore in my Understanding Intelligent Agency reading list.
I found this presentation by Suspended Reason on the history of sociology interesting, though I’m somewhat skeptical of the claims made in it.
In a different chapter Gintis discusses the “mixing problem”, an interesting idea I hadn’t encountered before. To give an example: the Nash equilibrium strategy in rock-paper-scissors is to play each with ⅓ probability. But given that your opponent is playing the equilibrium strategy, any strategy gets equal payoff for you. So the equilibrium strategy sometimes isn’t more rational than any other strategy, even conditional on your opponent’s strategy.
Infra-Bayesianism claims to have made progress in this direction, though I haven’t yet figured out how to get any clear insight from it.
I particularly liked Fink’s Clinical Introduction to Lacanian Psychoanalysis, which I found much more accessible than reading Lacan himself. I also appreciated Sadly, Porn and Existential Kink, which I think of as masculine- and feminine-flavored expositions of Lacanian ideas, respectively.
Another idea which is related (but not directly linked to the power theorists, to my knowledge) is the concept of status. Henrich and Gil-White distinguish between two kinds of status: prestige and dominance. Meanwhile Hanson and Simler describe the ubiquity of status-seeking behavior in modern life, including self-deceptive status-seeking. I haven’t included these in the main text because neither of these don’t do much to define or ground the concept of status in a rational agent setting.
Some central examples of work which engages with similar ideas, but doesn’t seem to be trying to build towards an elegant unified theory: Beyond Preferences in AI Alignment (Zhi-Xuan et al., 2024), A theory of appropriateness with applications to generative artificial intelligence (Leibo et al., 2024), and Full-stack alignment (Edelman et al., 2025). As one illustration: Leibo et al. fill in the hard part of their theory by simply asking an LLM to categorize appropriateness, which inspired me to post this meme.

