My Questions About Agency

One of the things I’ve spent a significant amount of time on is some Agent Foundations-adjacent philosophy work.

While I’ve put together reading lists and simulation environments, I have mostly engaged in drips and drabs to figure out what I will spend time thinking about when I get around to being focused on it directly. As a starting point for that, I think it’s worth laying out the sorts of questions I’m interested in and what sorts of answers would be satisfying, to keep me honest and focused as I dive into the literature of what has been worked out under which assumptions and constraints.

The first starting point is a generalization of the prisoner’s dilemma: what strategies can achieve Pareto-improvement over Nash equilibria? This generalizes in multiple directions with repeated games, number of players, etc. etc.; on a purely mathematical level I would like to know if there are classification theorems for the relevant topological objects.

A major source of unresolved problems in this space is the problem of externalities; in many arrangements between parties, there are parties not involved in the arrangement who also have an interest in the arrangement. Carbon emissions and climate impact are an illustrative example, where agreements (e.g. sale of gasoline for cars) generally have enough economic gain that taxes could be imposed on sales that could compensate the rest of humanity for pareto gains, but the external party is highly diffuse and may have difficulty estimating the impact or enforcing their interests in agreements.

The general version of this question is, how can one determine when an external party has an interest in a deal? In full generality there are obviously adversarial examples (e.g. an external party whose utility function is ~the inverse of one of the initial parties) that make the problem unsolvable. In practice, humans learn to identify (at least some) other humans as agents quite reliably, and at least somewhat consistently to assign partial agency to different animals. There is mathematical machinery underlying this for identifying “how goal directed” an actor is, or extracting goal-directed systems from causal diagrams. When exactly is this computationally tractable? What assumptions are required?

This machinery probably (?) mostly (?) generalizes to profinite limits, which is a necessary condition for the technically-finite but computationally-intractable level of detail in human perceptions to be able to implement these computational strategies. In practice humans shortcut much of the actor-identification problem by hard-wiring basic facial recognition, which leaves the problem of generalizing to non-human actors much more open. How else can we ground finite or profinite actor candidate sets?

The grounding analogy of biological similarity of humans generalizes with limited complication across most mammals, as well as to other animals with strong perceptive machinery (birds, cephalopods). This also motivates questions about computational complexity, perceptive capacity, and world-modeling capacity–not only are our perceptions finite, needing to converge to profinite limits, but our selves are also finite, needing to approximate profinite limits with some limited precision. What the nature of these limits is and how they affect agency is an open question strictly broader than the complex open questions of animal intelligence research.

Another direction speculated about in science fiction is collective intelligence, concerning the behavior of groups of animals, especially eusocial insects like ant and bee colonies. How does the mathematical machinery of agency adapt to this situation? What would it mean to negotiate a mutually beneficial agreement with an ant colony? Could a forest or an ecosystem act as a distributed agent? While my understanding is that mycology networks are not strong candidates for this being true in our world, could it work in theory?

The present day contains a new, non-biological examples of potential agents, in the form of LLMs. While a “base model” LLM is just a computer program with signature (string -> string), which matches none of our machinery, agentic harnesses “act in the world” in some sense, and could therefore display some of the types of behavior our machinery would classify as “agent-like”. Do they actually display this behavior? Under which circumstances? Does that create a moral obligation on users? Since the biological mechanisms (perception -> nervous system -> movement, need to eat and breathe) that make preferences grounded in animals don’t exist, a fully behavioral-mathematical explanation is needed to make progress.

Another open question is whether there are other actor candidates in our environment that we are missing. While intractable in full generality, and while the true fundamental nature of the universe is not known [citation needed], day to day we make do with Newtonian mechanics in Euclidean space [citation needed], which gives us a semi-natural substrate of objects and distance for grouping perceptions into a profinite set of candidates for actors and actions. This sort of pattern matching may relate to early human cultural beliefs in Animism, and I am especially curious if there is an agent-adjacent category (i.e. fulfilling some criteria but not others) that does naturally cover systems such as weather.

Regarding the known limits of Newtonian physics, relativity impacts our lives in tangible ways and introduces complexity for interstellar agents, and this divergence has a foundation of speculative fiction. I’m not aware of significant speculation on subatomic agents in which quantum mechanical failures of our notions of “object” and “distance” would be relevant, and I expect this direction to be more mathematical curiosity than active search-for-sentience, though I would not fully dismiss it out of hand.

Next
Next

LLM-driven Code Review Process