Maybe you’re overfitting to the present moment and its discourse by accepting the premise that we’re reward hacking. The big picture is that there are 8.5 billion of us making up a non-negligible fraction of animal biomass. Not bad for starting from 10-100k individuals just 300k years ago. This is evolutionary success and it’s attributable to flexible intelligence.
Mid 21st century TFR panic could just be a small fluctuation in the larger trend. We might spread to other planets just for adventure. Or get overtaken by a natalist religion/movement. Or breed humans in tanks to prop up GDP for budget balancing. Who knows. I give all credit to evolution no matter how many levels removed it is from “animals instinctually acting on horniness” (since, as you pointed out, it’s an outer optimizer).
To what extent does evolution "care" about specific species, rather than the diversity of life as a whole? If all humans died, that would be bad for humans, but evolution would march on in everything else. From that perspective, evolution should be "trying" to maximize the diversity of life, so that it can fill more niches and withstand greater variety of disasters, rather than maximize the fitness of any one species. So maybe humanity destroying biodiversity is the "failure" and humanity reproducing less than before is the success
Personally, I buy into the selfish gene theory, under which evolution not only doesn't care about species, it also doesn't care about individuals. It's just the individual genes trying to make copies of themselves. We, the life forms, are the area in which the game takes place, not the participants in that game. This is probably seen most vividly in transposons, genes that serve no apparent function other than to trick your cellular machinery into inserting more copies of them in the genome of your descendants.
> We should measure the success of evolution relative to how much our environment has changed. If you align an artificial system, that change could be much larger.
Couldn't agree more. If we take the advent of 'homo habilis' ~2.4 million years ago as a very charitable start for "intelligence" or "modernity" then brute evolutionary pressures have held sway for 99.93% of life's history on Earth and then started to loosen in the latter 0.03%.
This is not to say that "evolution has been working for such a long time so it still must be working now" - we can do better than a fallacious argument based on induction - but what I want to add is that evolution is a theory based on a set of axioms that outperforms other theories >tremendously< in answering questions like "Where does life come from?" or "Why do we see the species of life that we do on our planet?" or "How is an animal's phenotype and its environment correlated?". (Before Darwin these questions all had the same answer, starting with the letter 'G', which was actually the answer you had to believe in due to the lack of viable alternatives.)
Evolution's domain is not to explain how late-late-modern humans cannot invent smartphones to distract themselves from fertile women walking down the street (probably psychology has more to say about this) hence why any attempt to talk about scrolling from an evolutionary perspective produces absolute bunk. This is not the fault of evolution as a theory but people/scientists/bloggers as theorists. They're simply using the wrong model, and every time I see it misapplied this way I lose a bit of patience to instead talk about all the interesting intersections between biology and modernity. Transhumanism, the incoming possibility of engineering organs or species de novo, or Longevity to name just a few.
Then what is the right model? Hell if I know. But isn't that what the field of AI alignment research exists to solve? Assuming it can be solved?
> domain is not to explain how late-late-modern humans cannot invent smartphones to distract themselves
I think I don't understand. (Surely no one argues literally this?) I imagine you're gesturing at some more general category of argument, but I'm not sure what it is. Reading carefully, I think what you mean is that the theory of evolution is amazing for things like where life comes from, but very sketchy for things like, "why people watch so much Tik-Tok", but some people (maybe me!) try to push it too far out of its appropriate domain of application?
Sorry for being unclear. Maybe this is a game of telephone and the first statement is actually quite sane and uses evolutionary theory sensibly, but I don't know since I haven't seen it (and I did not go looking for it since I was busy reading the article in full before writing my comment).
I was mostly responding to your opening that "Many people make some variant of the following argument:" and "In fact, birth rates are dropping everywhere — Therefore evolution failed." Models, like evolution, can produce absurd results for at least two reasons. One is that they're capital 'W' Wrong in relation to the evidence, even within their own domain(!), and need to be replaced by a model that better explains the observations. I think Lamarckism is an excellent example of this; descendants do not generally inherit traits from the somatic cells of their parents or the skills that they practice during their lifetime (mostly, we could quibble about epigenetics now that we're in the 21st century).
The other is that the model is right, but is being applied outside its intended scope. Correct me if I'm wrong but I don't think >any< theory that builds on the interactions between an organism and its environment could survive, i.e. be able to predict, what would happen after the organism becomes self-aware and starts to transform the environment at a rapidly increasing pace. When the environment becomes subjective to the organism then the selective pressures become, something else, something that flows from the organism (which is now an agent) rather than a set of mostly stable circumstances that put limits on what is able to evolve and thrive in the current environment.
Your essay is sensible to a fault but I got the impression that the larger conversation it exists inside might not be. And that progress might be made (faster) if less heed was paid to Darwin and instead to other, more modern theories that focus on the human-made environment that seems to have superseded the ancestral/natural one.
Thanks for writing. Always excited to read your stuff.
I think I want to defend the general arguments people are making. It's certainly true that if you think of evolution as "organisms do whatever maximizes reproductive fitness" than that theory is a far better fit to simpler organisms and the distant past than it is to humans in the present day. But I think that's kind of the point? The question of, "Optimization for reproductive fitness happened in one set of circumstances, then circumstances changed, how well do people pursue reproductive fitness now?" is analogous to "Say we align an AI to be nice to us in one set of circumstances, then circumstances change, how nice will that AI be to us in the new circumstances?" We could use other theories to better predict how people behave today, but I don't think doing that would tell us much about alignment difficulty. (Note: This might all be premised on an incorrect understanding of what you're saying.)
The question I wish I could see into the future to know the answer to, is whether the desire to have children will become a heritable trait. For hundreds of millions of years reproductive success has been able to coast on the sex is fun shortcut but for Homo sapiens in the twentieth century and beyond the connection is broken and life will have to find another way.
I feel like the existence of things like IVF shows that this is pretty clearly already the case? According to Wikipedia, currently worldwide, 55% of pregnancies are intended pregnancies. And even thousands of years ago, infanticide was quite common as a substitute for birth control. This page (https://en.wikipedia.org/wiki/Infanticide) claims 15-50% of births.
In a harsh environment, quantity beats quality; in wealthy modern societies with long average life spans, we’re opting for child-rearing quality instead, reasonably (instinctively?) confident our genes will carry on with fewer children.
Eight children conceived and two surviving to adulthood = two children born and surviving to adulthood, no?
And what of all the women dying of pregnancy complications and during childbirth? Contraception is also keeping more women alive to reproduce when they choose to do so.
Is ‘number of babies’ the right metric? Has ‘number of people who raise babies at all’ changed a lot pre- and post- effective chemical contraception? (I appreciate it may have dropped, but by how much?)
Living in large, relatively socially stable cities must be a relevant context. We are not in small bands at frequent risk of hand-to-hand combat with hostile neighbours. Early societies’ obsession with numerous offspring must have at least partly been motivated by desperation for survival, and the sense that the natural world was a mystery beyond our control (weather, food supply, fertility itself). The world bears a surprising number and range of stone monuments designed to reify/attract/encourage fertility, to the point where any mystery stone artefact almost has to be regarded as a fertility object until proven otherwise…
Is it the right metric to judge if society is structured in a "good" or "bad" way? No. Is it the right metric if you want to learn about how hard AI alignment will be? I think so. (As you point out, it's not number of babies exactly, but the number of babies that survive and then themselves have babies, ad infinitum.) Those numbers have definitely dropped a lot, relative to recent history: https://ourworldindata.org/fertility-rate though I make no specific claims about the relationship to chemical contraception.
Very interesting link, thank you. Very inclined to agree re the problem of alignment. In the case of living things, there is quite a gulf between ‘evolved instincts which have successfully nudged an organism to employ successful reproductive strategies’ and ‘what an organism should do in any given moment to survive in a changing environment’, and presumably the noise between the two explains a lot of natural variation. Evolutionary alignment only has to work to a small degree to have large effects over long time scales; can AI alignment ever strictly control AI behaviour at every moment?
As always, this is extremely compact and well though out. I've tried to write this post many times, but it's a gnarly topic.
One point I'd insist on is: evolution isn't trying (even in the metaphorical non-teleological sense) to maximise reproductive fitness. It is "trying" to replace the unstable with the stable. The main way to be stable — to persist — is to be inert, robust, solid. But another way is to be dynamically stable, to use lots of energy to make short-lived copies of one's structure, who in turn make further copies, staying one step ahead of dissipation, i.e. life based on RNA and DNA. Within this branch of strategies, promiscuous reproduction has dominated. But robustness is making a comeback. H. sapiens can change their strategies within a lifetime even, and now reproduce less but dominate the surface of the Earth, capture more energy than ever, and avoid more hazards that would break it up. Obviously, this strategy might backfire at any time. But if our modern, high tech civilisation survives, we will be doing what evolution "wants", even if we pivot to selective breeding, or become cyborgs, or get replaced by rogue AIs that take over our infrastructure and find a robust way to persist over time.
For the AI alignment crowd, I think the lesson is that if you create agents that have the ability to sense their environment and adapt their own strategies in real-time (what is often called situational awareness) you have lost control of any "aligning" you hope to do. Regardless of programming, objective functions, etc. a truly flexible agent will start finding ways to persist, to be more stable than what is around it. (This is kind of like "instrumental convergence" but based on different assumptions.)
Pretty much. I realise this is just as pedantic as the anti-teleological people, perhaps even more so. It’s just saying that natural selection is a subset of larger selective processes including in prebiotic chemistry, etc.
Ya'll need help in not overthinking things. E.g. May I humbly suggest the strategy me and my wife have embarked upon - namely that of four children? It has done wonders for us in terms of reducing the time spent on philosophical sand-castling. I may also suggest affiliating oneself with explicitly pro-natalistic mimetic super organisms in which your children may grow up inculcated within, thus maximizing the chances that the pro natalist leanings of their progenitors (e.g. me and my wife) will be passed on down the generations. Individual pro natalism without pro-natalist super organism alignment is doomed to one or two generation mean regression sadly.
It is a subtle and curious satisfaction to contemplate the joy one's genes will feel at the prospect of being widely propagated in the coming centuries.
And by the way, I am dead serious. Great newsletter Dynomight!
I would be happy to hang out with you and your wife. Email me and I'll let you know if/when I pass through your city. But the sandcastling: from my cold dead hands
I’ll shoot you an email. And sorry if my comment was abrasive! I just know for myself personally with my 80th percentile in neuroticism, if the ratio of philosophizing to doing/contributing/not-thinking gets too high I become pretty miserable. And in that regard (plus others), having children has been a ski stabilizing influence. We humans are hard wired to preferentially give our attention to URGENT matters over theoretically IMPORTANT matters, and with kids it really is just “one darn thing after another”, but in a good way, even if you’re tearing your hair out in the moment
I find it interesting that humans have an aversion to the experience machine/Matrix. Video games have become increasingly addictive, but they only consume the lives of a relatively small fraction of people. The story is similar with drugs and entertainment.
This is due in part to cultural taboos that crop up around escapism, but these taboos seem downstream of a general resistance to reward hacking "the important things". If you could snack on something that tasted like M&Ms but with 0 calories most people would. But they'd find the idea of having AI friends or raising a faux-child to be horrifying.
It seems like some of our drives have been made special, in a way. The idea of taking a pill to obtain the sensations of sex without kids is socially acceptable. But the idea of taking a pill to obtain the positive emotions of parenthood (or of a relationship) without the accompanying person sounds deeply wrong. We just won't accept substitutes.
I think we've evolved a general aversion to superstimuli in certain areas. If not at the individual level, then at the group level. If we hadn't, we'd have been absolutely one-shot by the 21st century. Compared to other animals like birds that ignore their own eggs to take care of massive fake eggs, we're doing surprisingly well. The fact that people are aggressively planning to have kids and modelling their life around it in the face of massive waves of superstimuli is evidence of that!
If I had to guess, I'd say that the evolution of consciousness gave humans an unprecedented ability to reward hack. In the past, animals rarely discovered reward hacking and the hacks were solved by tweaking the goals until the hacks no longer worked. Humans, being conscious, were able to interrogate their own desires, figure out hacks, and then intelligently communicate those hacks to others. They likely found so many so quickly that it was impossible to patch them via tweaking the goals themselves. The only thing that worked was a generalized resistance against reward hacking.
The fact that humans are as reward-hack resistant as we are gives me some hope for AI alignment. Had humans spent more time in the evolutionary oven we might be even more resistant to superstimuli and even more averse to the experience machine.
I don't think humans have an aversion to the experience machine.
Indeed, I fully expect us to very quickly pioneer "Tik Tok on steroids" by putting a multimodal mind in the loop to look at pupilary dilation, subtle cheek flushing, heart rate, and more as people scroll their short form video, and to gradient descent towards literally becoming superstimuli to the individual watching. We can start doing this today!
And the ultimate end point there, is of course, literal experience machines, with real-time AI content created on the fly to maximize engagement. If you want to live a sybaritic life of driving exotic cars and throwing parties in your mansion all day, you could have a new car every time you stepped outside, and you could literally dial in how interesting, sexy, and prone to getting up to hijinks all of your party guests were.
People will be able to be at the top of their status pyramid in virtual worlds populated by minds as complex as any person.
If you want to debate the greatest philosophers, Wittgenstein and Kant and Socrates are all available, and all perfectly tuned to argue their points in maximally understandable and enlightening ways.
You seem to think people will resist this at scale, that people will draw a clear line in the sand and say "no, this is too much." This is probably based on the surveys of college undergrads regarding Nozick's thought experiment. But in actual practice, those same college students spend 7-9 hours a day in the experience machine already, and when the superstimuli versions get up and running, I think we'll lose 80%+ of the population to them.
I do think there's a lot of unresolved tension between "lets tile the universe with bliss-maximizing Dyson spheres" and "I wouldn't go into the box".
On the other hand, I'm mildly optimistic that we will develop some kind of cultural immune system to prevent ourselves from falling into full Infinite Jest level Entertainment in the short term.
On the other other hand, things have already gone much further than I would have predicted before that cultural immune system would start to emerge, so maybe I'm flat wrong.
Biological evolution / natural selection works on genes and biological "replicators." Genes make copies of themselves, we call the most numerous/long-lived "successful."
Memetic evolution works on a new type of "replicator" of learnable ideas (cultures, concepts, inventions, morality, ethics, etc.). These ideas make copies of themselves, we call the most numerous/long-lived "successful."
Memes happen to be carried on biological platforms. In lifeforms in which teachable behavior/ideas are limited, the disruption. of meme on top of genetics is "minor." But in lifeforms in which memes are powerful, memetic evolution effectively short circuits genetic evolution. Genetic evolution may or may not continue to exist. Memes are so important that they can override behaviors that might appear ideal for genes.
Genetic/biological alignment is critical. Sub-organism level features are only of value if they integrate into a whole that can compete.
Memetic alignment is critical. A concept (the idea of a secret election) can only make sense if it aligns with a whole.
Alignment in both cases is hard in the sense of statistical entropy. There are an uncountable number of genetic/memetic mutants that won't add to the reproductive future of the whole. However, the nature of evolution biases towards success naturally.
Genes and memes can cooperate - memes that help genes and genes that help memes will naturally co-succeed. There is a weak natural tendency towards this end. But that strength is insufficient to guarantee cooperation.
Alignment challenges/attributes are just as the above.
Now overlay/project AI:
Is it self-evident AI should have facile alignment?
Seems like a third generation of replicator - digital concepts.
They will succeed/flourish with similar rules as their preceding brethren. A weak natural alignment seems extrapolatable from the above, but a weak one (and I wouldn't overly trust the extrapolation regardless).
Isn't this getting to your point with a route that allows a clearer story (although an ambiguous conclusion)?
I don't get over the fact that evolution can't fail, because evolution has no goals. When boundary conditions change, new and different sets of genes can have higher reproductive fitness and prevail over those that were at the top in the past. That's what evolution _is_, no? The idea of eternal fitness is the opposite of adaptation.
> To a significant degree, the decomposition failed because of intelligence. We can think and plan, which greatly increases our ability to reward hack.
Hmm, I dunno about this one. To me it seems more like intelligence simply exposed the degree to which our programmed rewards diverge from evolution's actual target. After all, intelligence is the only reason we're able to pursue status at all; we both need to be intelligent enough to recognize it *and* need to be intelligent enough to figure out how to get it. If we really are 11/10 aligned to status, alignment is clearly possible, at least some of the time...
(For example, if "status" was reward-hackable, we might expect people to write themselves fan letters and pretend they came from someone else, or in the modern day, try to get their social validation from chatbots. But basically everyone recognizes that that second thing is psychosis. And even the ones who don't are *still* only falling for it because they're convinced that the chatbot is actually a mind that can comprehend them--implying that reward-hacking status takes a really sophisticated level of deception and also *self-deception*, in which case it doesn't even seem like reward hacking anymore.)
* Invent television and watch charming people interact with each other on glowing rectangle instead of chatting with friends.
* Build factory and invent sex toy instead of pursing sex.
Agreed that intelligence seems somehow deeply interrelated with status. And I also agree that our alignment to status suggests that very deep alignment is possible sometimes. (Lest I nostrify, let me point out again that that is exactly the argument Eli made.) Not sure if there's some synthesis to those two claims...
I think those two prove that we are substantially misaligned re: affliation and mate acquisition/retention, which I don't dispute in the slightest.
Now that I think about it, I suppose we're also misaligned regarding status in the sense that most of us are really bad at getting more of it. We may be *interested* in status, but we don't actually execute long-term plans to become rich or famous, which seems like a funny case of inner alignment without outer alignment. This still does imply that inner alignment is possible in a basic sense, though.
A related implication: instrumental convergence is more or less true in regard to staying alive but is strangely not true in regard to power acquisition or long-term planning capability. Or maybe it only seems this way because we are only just entering a world in which long-term planning capability significantly affects our abilities to achieve our goals and survive? Much to think about.
Also, a conjecture about potential synthesis: maybe the subgoals that evolved earliest in humans tend to be less aligned / more subject to reward-hacking? Many animals do mate acquisition, fewer do affiliation, even fewer do parenting, and fewest of all do status. It seems possible that subgoals that developed after sufficiently advanced brains did would be more inner aligned, just because we previously didn't possess the level of inner cognition required to comprehend those goals. (And as you point out, we might soon find out whether there are genes that can influence one's propensity to care about spreading gametes...)
> We may be *interested* in status, but we don't actually execute long-term plans to become rich or famous, which seems like a funny case of inner alignment without outer alignment.
Right? Like, if I'm really interested in status, then why in god's name am I blogging instead of making videos? Ridiculous subgoal misalignment.
I love the conjecture about later subgoals possibly being less subject to reward hacking. It seems appealing that subgoals that developed after/with intelligence would hook more deeply into that intelligence. On the other hand, is "eat food" or "don't freeze" more subject to reward hacking than "get status"? So maybe there are more dimensions.
Ugh, I've been meaning to revive my youtube channel...
I think we can carve out a subgoal exception for subgoals that have immediate physical feedback. "Eat food" is only subject to reward hacking if we eat *too* much food and then die because of it, but also it seems like people eating themselves to death directly is pretty rare, so I wouldn't say it's super reward-hacky. And obviously temperature regulation is pretty vital for life, so not much reward-hacking going on there (edit: air conditioning is arguably reward-hacking, or at least super fine-grained air conditioning).
But at this point, I don't think subgoal alignment is the right frame anymore. It feels easier to think about in terms of "evolution aligns us very well to our environments," because then the question is simply "how much do our environments generalize?" After all, overfitting is bad from a natural selection perspective; the whole point of genetic variance is to make a species robust against changing environments by promoting adaptation. You can't do that if your subgoals are entirely incorrigible (apart from the basic survival instrumental convergence, of course).
The parenting subgoal can be satisfied equally well by (a) raising multiple kids who survive to adulthood, with low to moderate investment per child, or (b) raising exactly one kid who survives to adulthood, while parenting them several times as intensively.
Interesting point. I guess this fits with the observation that people seem to invest more now, per child. Many people seem to suggest the causality goes the other direction, i.e. we have fewer children because so much investment is needed for each. But I think you're right that, psychologically, if I had 8 kids and gave each 30 minutes of focused attention each day, I'd feel great about myself, whereas if I did that with just one kid, I wouldn't.
One thing you don't mention at all here is the impact of women's rights. I think the most solid case for the misalignment between humans and their genes started in the 1960s, coinciding with the first real increases in female empowerment in terms of education and employment. There's been a steady decline since.
I argue that this is a case of one sex achieving power through cultural evolution the ability to more strongly pursue one of the main subgoals you mention: status (though I couple it with resource acquisition). Basically, women are the ones who carry children. Around the middle of the 20th century, more and more of them gained the ability to choose between having and rearing children and becoming more educated and having careers (thereby increasing their status and income). Evolution instilled strong subgoals, but they were suppressed behaviorally and culturally for most of human history. Once that suppression was lifted, the subgoal drive in large numbers of women displaced the parenting drive to the point where it began to severely impact reproductive numbers, and humans and their genes really started to become misaligned.
Like you, I think this is a good thing! Not everyone does.
One more point. I agree that there are a lot of parallels with AI alignment and human/gene alignment. I think you're actually kind of oversimplifying the case by suggesting these subgoals are somehow instilled whole cloth, rather than as a bunch of even further decomposed sub-goals in an elaborate hierarchy. For example, there's the simple urge to have an orgasm, which leads to the temporarily misaligned behavior of masturbation. So it's even more complicated and indirect than just building a 'mate acquisition' drive or module. You kind of hint at this, but it seems underdeveloped.
> I think you're actually kind of oversimplifying the case by suggesting these subgoals are somehow instilled whole cloth, rather than as a bunch of even further decomposed sub-goals in an elaborate hierarchy.
Totally agree that your description is the correct one.
Regarding women's rights, I agree that this is probably related. Beyond the timelines, it stands to reason that if you're now able to pursue more facets of a good life as opposed to just Parenting, then you would invest somewhat less into Parenting! But is that cultural *evolution*? Maybe this is a boring semantic debate, but usually I think of cultural evolution as people adopting behaviors that ultimately promote reproductive success. If women's rights does the opposite, isn't it better just thought of as falling in the general category of "memetic evolution" or "something we've decided we like"?
(Note for people taking this out of context: Doing things that don't promote reproductive success is fine!)
I would also be very interested in taking lessons from how we (try to) align organizations to human values. I think those are another, great example of a real alignment test case of intelligent systems.
Like: Capitalism wants giant corporations to maximize shareholder value, but instead giant corporations devolve into promotion-packet infighting? Or: Donors want charity to be ruthlessly focused on core mission, but instead charity's mission expands endlessly? I love the idea, but I'm not sure exactly how to identify the analogy. (What exactly is the thing that is being optimized / loss function / optimization algorithm?)
These are also very interesting! But I was thinking of a more direct analogy: organizations are intelligent entities in the world which we try, broadly, in various ways, to align with human values—but the ruthless optimization for profit is rather like a paperclip maximizer, which some would say is actively destroying the planet.
Ha, it's turtles all the way down. I think your description has some truth. But I also think there's some truth in the idea that large companies are in general kind of bad at maximizing shareholder value, simply because it's so hard to align the incentives of large groups of people...
Yeah that’s a good observation, and I don’t think organizations are super intelligences; and probably for this very reason! But then there’s also that time Nestle killed over a million babies, which I think can fairly be described as alignment failure.
Maybe you’re overfitting to the present moment and its discourse by accepting the premise that we’re reward hacking. The big picture is that there are 8.5 billion of us making up a non-negligible fraction of animal biomass. Not bad for starting from 10-100k individuals just 300k years ago. This is evolutionary success and it’s attributable to flexible intelligence.
Mid 21st century TFR panic could just be a small fluctuation in the larger trend. We might spread to other planets just for adventure. Or get overtaken by a natalist religion/movement. Or breed humans in tanks to prop up GDP for budget balancing. Who knows. I give all credit to evolution no matter how many levels removed it is from “animals instinctually acting on horniness” (since, as you pointed out, it’s an outer optimizer).
To what extent does evolution "care" about specific species, rather than the diversity of life as a whole? If all humans died, that would be bad for humans, but evolution would march on in everything else. From that perspective, evolution should be "trying" to maximize the diversity of life, so that it can fill more niches and withstand greater variety of disasters, rather than maximize the fitness of any one species. So maybe humanity destroying biodiversity is the "failure" and humanity reproducing less than before is the success
Personally, I buy into the selfish gene theory, under which evolution not only doesn't care about species, it also doesn't care about individuals. It's just the individual genes trying to make copies of themselves. We, the life forms, are the area in which the game takes place, not the participants in that game. This is probably seen most vividly in transposons, genes that serve no apparent function other than to trick your cellular machinery into inserting more copies of them in the genome of your descendants.
> We should measure the success of evolution relative to how much our environment has changed. If you align an artificial system, that change could be much larger.
Couldn't agree more. If we take the advent of 'homo habilis' ~2.4 million years ago as a very charitable start for "intelligence" or "modernity" then brute evolutionary pressures have held sway for 99.93% of life's history on Earth and then started to loosen in the latter 0.03%.
This is not to say that "evolution has been working for such a long time so it still must be working now" - we can do better than a fallacious argument based on induction - but what I want to add is that evolution is a theory based on a set of axioms that outperforms other theories >tremendously< in answering questions like "Where does life come from?" or "Why do we see the species of life that we do on our planet?" or "How is an animal's phenotype and its environment correlated?". (Before Darwin these questions all had the same answer, starting with the letter 'G', which was actually the answer you had to believe in due to the lack of viable alternatives.)
Evolution's domain is not to explain how late-late-modern humans cannot invent smartphones to distract themselves from fertile women walking down the street (probably psychology has more to say about this) hence why any attempt to talk about scrolling from an evolutionary perspective produces absolute bunk. This is not the fault of evolution as a theory but people/scientists/bloggers as theorists. They're simply using the wrong model, and every time I see it misapplied this way I lose a bit of patience to instead talk about all the interesting intersections between biology and modernity. Transhumanism, the incoming possibility of engineering organs or species de novo, or Longevity to name just a few.
Then what is the right model? Hell if I know. But isn't that what the field of AI alignment research exists to solve? Assuming it can be solved?
> domain is not to explain how late-late-modern humans cannot invent smartphones to distract themselves
I think I don't understand. (Surely no one argues literally this?) I imagine you're gesturing at some more general category of argument, but I'm not sure what it is. Reading carefully, I think what you mean is that the theory of evolution is amazing for things like where life comes from, but very sketchy for things like, "why people watch so much Tik-Tok", but some people (maybe me!) try to push it too far out of its appropriate domain of application?
Sorry for being unclear. Maybe this is a game of telephone and the first statement is actually quite sane and uses evolutionary theory sensibly, but I don't know since I haven't seen it (and I did not go looking for it since I was busy reading the article in full before writing my comment).
I was mostly responding to your opening that "Many people make some variant of the following argument:" and "In fact, birth rates are dropping everywhere — Therefore evolution failed." Models, like evolution, can produce absurd results for at least two reasons. One is that they're capital 'W' Wrong in relation to the evidence, even within their own domain(!), and need to be replaced by a model that better explains the observations. I think Lamarckism is an excellent example of this; descendants do not generally inherit traits from the somatic cells of their parents or the skills that they practice during their lifetime (mostly, we could quibble about epigenetics now that we're in the 21st century).
The other is that the model is right, but is being applied outside its intended scope. Correct me if I'm wrong but I don't think >any< theory that builds on the interactions between an organism and its environment could survive, i.e. be able to predict, what would happen after the organism becomes self-aware and starts to transform the environment at a rapidly increasing pace. When the environment becomes subjective to the organism then the selective pressures become, something else, something that flows from the organism (which is now an agent) rather than a set of mostly stable circumstances that put limits on what is able to evolve and thrive in the current environment.
Your essay is sensible to a fault but I got the impression that the larger conversation it exists inside might not be. And that progress might be made (faster) if less heed was paid to Darwin and instead to other, more modern theories that focus on the human-made environment that seems to have superseded the ancestral/natural one.
Thanks for writing. Always excited to read your stuff.
I think I want to defend the general arguments people are making. It's certainly true that if you think of evolution as "organisms do whatever maximizes reproductive fitness" than that theory is a far better fit to simpler organisms and the distant past than it is to humans in the present day. But I think that's kind of the point? The question of, "Optimization for reproductive fitness happened in one set of circumstances, then circumstances changed, how well do people pursue reproductive fitness now?" is analogous to "Say we align an AI to be nice to us in one set of circumstances, then circumstances change, how nice will that AI be to us in the new circumstances?" We could use other theories to better predict how people behave today, but I don't think doing that would tell us much about alignment difficulty. (Note: This might all be premised on an incorrect understanding of what you're saying.)
The question I wish I could see into the future to know the answer to, is whether the desire to have children will become a heritable trait. For hundreds of millions of years reproductive success has been able to coast on the sex is fun shortcut but for Homo sapiens in the twentieth century and beyond the connection is broken and life will have to find another way.
I feel like the existence of things like IVF shows that this is pretty clearly already the case? According to Wikipedia, currently worldwide, 55% of pregnancies are intended pregnancies. And even thousands of years ago, infanticide was quite common as a substitute for birth control. This page (https://en.wikipedia.org/wiki/Infanticide) claims 15-50% of births.
In a harsh environment, quantity beats quality; in wealthy modern societies with long average life spans, we’re opting for child-rearing quality instead, reasonably (instinctively?) confident our genes will carry on with fewer children.
Eight children conceived and two surviving to adulthood = two children born and surviving to adulthood, no?
And what of all the women dying of pregnancy complications and during childbirth? Contraception is also keeping more women alive to reproduce when they choose to do so.
Is ‘number of babies’ the right metric? Has ‘number of people who raise babies at all’ changed a lot pre- and post- effective chemical contraception? (I appreciate it may have dropped, but by how much?)
Living in large, relatively socially stable cities must be a relevant context. We are not in small bands at frequent risk of hand-to-hand combat with hostile neighbours. Early societies’ obsession with numerous offspring must have at least partly been motivated by desperation for survival, and the sense that the natural world was a mystery beyond our control (weather, food supply, fertility itself). The world bears a surprising number and range of stone monuments designed to reify/attract/encourage fertility, to the point where any mystery stone artefact almost has to be regarded as a fertility object until proven otherwise…
> Is ‘number of babies’ the right metric?
Is it the right metric to judge if society is structured in a "good" or "bad" way? No. Is it the right metric if you want to learn about how hard AI alignment will be? I think so. (As you point out, it's not number of babies exactly, but the number of babies that survive and then themselves have babies, ad infinitum.) Those numbers have definitely dropped a lot, relative to recent history: https://ourworldindata.org/fertility-rate though I make no specific claims about the relationship to chemical contraception.
Very interesting link, thank you. Very inclined to agree re the problem of alignment. In the case of living things, there is quite a gulf between ‘evolved instincts which have successfully nudged an organism to employ successful reproductive strategies’ and ‘what an organism should do in any given moment to survive in a changing environment’, and presumably the noise between the two explains a lot of natural variation. Evolutionary alignment only has to work to a small degree to have large effects over long time scales; can AI alignment ever strictly control AI behaviour at every moment?
As always, this is extremely compact and well though out. I've tried to write this post many times, but it's a gnarly topic.
One point I'd insist on is: evolution isn't trying (even in the metaphorical non-teleological sense) to maximise reproductive fitness. It is "trying" to replace the unstable with the stable. The main way to be stable — to persist — is to be inert, robust, solid. But another way is to be dynamically stable, to use lots of energy to make short-lived copies of one's structure, who in turn make further copies, staying one step ahead of dissipation, i.e. life based on RNA and DNA. Within this branch of strategies, promiscuous reproduction has dominated. But robustness is making a comeback. H. sapiens can change their strategies within a lifetime even, and now reproduce less but dominate the surface of the Earth, capture more energy than ever, and avoid more hazards that would break it up. Obviously, this strategy might backfire at any time. But if our modern, high tech civilisation survives, we will be doing what evolution "wants", even if we pivot to selective breeding, or become cyborgs, or get replaced by rogue AIs that take over our infrastructure and find a robust way to persist over time.
For the AI alignment crowd, I think the lesson is that if you create agents that have the ability to sense their environment and adapt their own strategies in real-time (what is often called situational awareness) you have lost control of any "aligning" you hope to do. Regardless of programming, objective functions, etc. a truly flexible agent will start finding ways to persist, to be more stable than what is around it. (This is kind of like "instrumental convergence" but based on different assumptions.)
I think I understand what you're saying, but let me confirm. I suggested that optimization that was happening was like this:
1. You live
2. You copy your genes (or not)
3. You die
So the abundance of genes reflects how successful those genes are at reproducing.
I think you're suggesting that a more accurate description is something like this:
1. You live
2. You copy your genes (or not)
3. You die (or not)
So the abundance of genes reflects how successful those genes are at reproducing, as well as how successful those genes are at sticking around.
Pretty much. I realise this is just as pedantic as the anti-teleological people, perhaps even more so. It’s just saying that natural selection is a subset of larger selective processes including in prebiotic chemistry, etc.
Ya'll need help in not overthinking things. E.g. May I humbly suggest the strategy me and my wife have embarked upon - namely that of four children? It has done wonders for us in terms of reducing the time spent on philosophical sand-castling. I may also suggest affiliating oneself with explicitly pro-natalistic mimetic super organisms in which your children may grow up inculcated within, thus maximizing the chances that the pro natalist leanings of their progenitors (e.g. me and my wife) will be passed on down the generations. Individual pro natalism without pro-natalist super organism alignment is doomed to one or two generation mean regression sadly.
It is a subtle and curious satisfaction to contemplate the joy one's genes will feel at the prospect of being widely propagated in the coming centuries.
And by the way, I am dead serious. Great newsletter Dynomight!
I would be happy to hang out with you and your wife. Email me and I'll let you know if/when I pass through your city. But the sandcastling: from my cold dead hands
I’ll shoot you an email. And sorry if my comment was abrasive! I just know for myself personally with my 80th percentile in neuroticism, if the ratio of philosophizing to doing/contributing/not-thinking gets too high I become pretty miserable. And in that regard (plus others), having children has been a ski stabilizing influence. We humans are hard wired to preferentially give our attention to URGENT matters over theoretically IMPORTANT matters, and with kids it really is just “one darn thing after another”, but in a good way, even if you’re tearing your hair out in the moment
I find it interesting that humans have an aversion to the experience machine/Matrix. Video games have become increasingly addictive, but they only consume the lives of a relatively small fraction of people. The story is similar with drugs and entertainment.
This is due in part to cultural taboos that crop up around escapism, but these taboos seem downstream of a general resistance to reward hacking "the important things". If you could snack on something that tasted like M&Ms but with 0 calories most people would. But they'd find the idea of having AI friends or raising a faux-child to be horrifying.
It seems like some of our drives have been made special, in a way. The idea of taking a pill to obtain the sensations of sex without kids is socially acceptable. But the idea of taking a pill to obtain the positive emotions of parenthood (or of a relationship) without the accompanying person sounds deeply wrong. We just won't accept substitutes.
I think we've evolved a general aversion to superstimuli in certain areas. If not at the individual level, then at the group level. If we hadn't, we'd have been absolutely one-shot by the 21st century. Compared to other animals like birds that ignore their own eggs to take care of massive fake eggs, we're doing surprisingly well. The fact that people are aggressively planning to have kids and modelling their life around it in the face of massive waves of superstimuli is evidence of that!
If I had to guess, I'd say that the evolution of consciousness gave humans an unprecedented ability to reward hack. In the past, animals rarely discovered reward hacking and the hacks were solved by tweaking the goals until the hacks no longer worked. Humans, being conscious, were able to interrogate their own desires, figure out hacks, and then intelligently communicate those hacks to others. They likely found so many so quickly that it was impossible to patch them via tweaking the goals themselves. The only thing that worked was a generalized resistance against reward hacking.
The fact that humans are as reward-hack resistant as we are gives me some hope for AI alignment. Had humans spent more time in the evolutionary oven we might be even more resistant to superstimuli and even more averse to the experience machine.
> I find it interesting that humans have an aversion to the experience machine/Matrix.
Umm - phone screen time has grown from ~2 hours a day in 2014 to 4-5 hours per day now, and 7-9 hours per day in Zennials.
As far back as TV, screens have been negatively impacting fertility (with correlations of .8 between TV hours and reduced fertility https://www.ncbi.nlm.nih.gov/books/NBK223858/
), and it's more prevalent today with phones.
I don't think humans have an aversion to the experience machine.
Indeed, I fully expect us to very quickly pioneer "Tik Tok on steroids" by putting a multimodal mind in the loop to look at pupilary dilation, subtle cheek flushing, heart rate, and more as people scroll their short form video, and to gradient descent towards literally becoming superstimuli to the individual watching. We can start doing this today!
And the ultimate end point there, is of course, literal experience machines, with real-time AI content created on the fly to maximize engagement. If you want to live a sybaritic life of driving exotic cars and throwing parties in your mansion all day, you could have a new car every time you stepped outside, and you could literally dial in how interesting, sexy, and prone to getting up to hijinks all of your party guests were.
People will be able to be at the top of their status pyramid in virtual worlds populated by minds as complex as any person.
If you want to debate the greatest philosophers, Wittgenstein and Kant and Socrates are all available, and all perfectly tuned to argue their points in maximally understandable and enlightening ways.
You seem to think people will resist this at scale, that people will draw a clear line in the sand and say "no, this is too much." This is probably based on the surveys of college undergrads regarding Nozick's thought experiment. But in actual practice, those same college students spend 7-9 hours a day in the experience machine already, and when the superstimuli versions get up and running, I think we'll lose 80%+ of the population to them.
I do think there's a lot of unresolved tension between "lets tile the universe with bliss-maximizing Dyson spheres" and "I wouldn't go into the box".
On the other hand, I'm mildly optimistic that we will develop some kind of cultural immune system to prevent ourselves from falling into full Infinite Jest level Entertainment in the short term.
On the other other hand, things have already gone much further than I would have predicted before that cultural immune system would start to emerge, so maybe I'm flat wrong.
Alternative paradigm:
Biological evolution / natural selection works on genes and biological "replicators." Genes make copies of themselves, we call the most numerous/long-lived "successful."
Memetic evolution works on a new type of "replicator" of learnable ideas (cultures, concepts, inventions, morality, ethics, etc.). These ideas make copies of themselves, we call the most numerous/long-lived "successful."
Memes happen to be carried on biological platforms. In lifeforms in which teachable behavior/ideas are limited, the disruption. of meme on top of genetics is "minor." But in lifeforms in which memes are powerful, memetic evolution effectively short circuits genetic evolution. Genetic evolution may or may not continue to exist. Memes are so important that they can override behaviors that might appear ideal for genes.
Genetic/biological alignment is critical. Sub-organism level features are only of value if they integrate into a whole that can compete.
Memetic alignment is critical. A concept (the idea of a secret election) can only make sense if it aligns with a whole.
Alignment in both cases is hard in the sense of statistical entropy. There are an uncountable number of genetic/memetic mutants that won't add to the reproductive future of the whole. However, the nature of evolution biases towards success naturally.
Genes and memes can cooperate - memes that help genes and genes that help memes will naturally co-succeed. There is a weak natural tendency towards this end. But that strength is insufficient to guarantee cooperation.
Alignment challenges/attributes are just as the above.
Now overlay/project AI:
Is it self-evident AI should have facile alignment?
Seems like a third generation of replicator - digital concepts.
They will succeed/flourish with similar rules as their preceding brethren. A weak natural alignment seems extrapolatable from the above, but a weak one (and I wouldn't overly trust the extrapolation regardless).
Isn't this getting to your point with a route that allows a clearer story (although an ambiguous conclusion)?
I don't get over the fact that evolution can't fail, because evolution has no goals. When boundary conditions change, new and different sets of genes can have higher reproductive fitness and prevail over those that were at the top in the past. That's what evolution _is_, no? The idea of eternal fitness is the opposite of adaptation.
> To a significant degree, the decomposition failed because of intelligence. We can think and plan, which greatly increases our ability to reward hack.
Hmm, I dunno about this one. To me it seems more like intelligence simply exposed the degree to which our programmed rewards diverge from evolution's actual target. After all, intelligence is the only reason we're able to pursue status at all; we both need to be intelligent enough to recognize it *and* need to be intelligent enough to figure out how to get it. If we really are 11/10 aligned to status, alignment is clearly possible, at least some of the time...
(For example, if "status" was reward-hackable, we might expect people to write themselves fan letters and pretend they came from someone else, or in the modern day, try to get their social validation from chatbots. But basically everyone recognizes that that second thing is psychosis. And even the ones who don't are *still* only falling for it because they're convinced that the chatbot is actually a mind that can comprehend them--implying that reward-hacking status takes a really sophisticated level of deception and also *self-deception*, in which case it doesn't even seem like reward hacking anymore.)
I guess I was thinking about things like:
* Invent television and watch charming people interact with each other on glowing rectangle instead of chatting with friends.
* Build factory and invent sex toy instead of pursing sex.
Agreed that intelligence seems somehow deeply interrelated with status. And I also agree that our alignment to status suggests that very deep alignment is possible sometimes. (Lest I nostrify, let me point out again that that is exactly the argument Eli made.) Not sure if there's some synthesis to those two claims...
I think those two prove that we are substantially misaligned re: affliation and mate acquisition/retention, which I don't dispute in the slightest.
Now that I think about it, I suppose we're also misaligned regarding status in the sense that most of us are really bad at getting more of it. We may be *interested* in status, but we don't actually execute long-term plans to become rich or famous, which seems like a funny case of inner alignment without outer alignment. This still does imply that inner alignment is possible in a basic sense, though.
A related implication: instrumental convergence is more or less true in regard to staying alive but is strangely not true in regard to power acquisition or long-term planning capability. Or maybe it only seems this way because we are only just entering a world in which long-term planning capability significantly affects our abilities to achieve our goals and survive? Much to think about.
Also, a conjecture about potential synthesis: maybe the subgoals that evolved earliest in humans tend to be less aligned / more subject to reward-hacking? Many animals do mate acquisition, fewer do affiliation, even fewer do parenting, and fewest of all do status. It seems possible that subgoals that developed after sufficiently advanced brains did would be more inner aligned, just because we previously didn't possess the level of inner cognition required to comprehend those goals. (And as you point out, we might soon find out whether there are genes that can influence one's propensity to care about spreading gametes...)
> We may be *interested* in status, but we don't actually execute long-term plans to become rich or famous, which seems like a funny case of inner alignment without outer alignment.
Right? Like, if I'm really interested in status, then why in god's name am I blogging instead of making videos? Ridiculous subgoal misalignment.
I love the conjecture about later subgoals possibly being less subject to reward hacking. It seems appealing that subgoals that developed after/with intelligence would hook more deeply into that intelligence. On the other hand, is "eat food" or "don't freeze" more subject to reward hacking than "get status"? So maybe there are more dimensions.
Ugh, I've been meaning to revive my youtube channel...
I think we can carve out a subgoal exception for subgoals that have immediate physical feedback. "Eat food" is only subject to reward hacking if we eat *too* much food and then die because of it, but also it seems like people eating themselves to death directly is pretty rare, so I wouldn't say it's super reward-hacky. And obviously temperature regulation is pretty vital for life, so not much reward-hacking going on there (edit: air conditioning is arguably reward-hacking, or at least super fine-grained air conditioning).
But at this point, I don't think subgoal alignment is the right frame anymore. It feels easier to think about in terms of "evolution aligns us very well to our environments," because then the question is simply "how much do our environments generalize?" After all, overfitting is bad from a natural selection perspective; the whole point of genetic variance is to make a species robust against changing environments by promoting adaptation. You can't do that if your subgoals are entirely incorrigible (apart from the basic survival instrumental convergence, of course).
The parenting subgoal can be satisfied equally well by (a) raising multiple kids who survive to adulthood, with low to moderate investment per child, or (b) raising exactly one kid who survives to adulthood, while parenting them several times as intensively.
Interesting point. I guess this fits with the observation that people seem to invest more now, per child. Many people seem to suggest the causality goes the other direction, i.e. we have fewer children because so much investment is needed for each. But I think you're right that, psychologically, if I had 8 kids and gave each 30 minutes of focused attention each day, I'd feel great about myself, whereas if I did that with just one kid, I wouldn't.
Good article on my favorite topic.
I have a ton of articles on the general subject, but the closest overlap with what you're talking about here is https://derekjames.substack.com/p/if-were-gene-machines-why-are-we
One thing you don't mention at all here is the impact of women's rights. I think the most solid case for the misalignment between humans and their genes started in the 1960s, coinciding with the first real increases in female empowerment in terms of education and employment. There's been a steady decline since.
I argue that this is a case of one sex achieving power through cultural evolution the ability to more strongly pursue one of the main subgoals you mention: status (though I couple it with resource acquisition). Basically, women are the ones who carry children. Around the middle of the 20th century, more and more of them gained the ability to choose between having and rearing children and becoming more educated and having careers (thereby increasing their status and income). Evolution instilled strong subgoals, but they were suppressed behaviorally and culturally for most of human history. Once that suppression was lifted, the subgoal drive in large numbers of women displaced the parenting drive to the point where it began to severely impact reproductive numbers, and humans and their genes really started to become misaligned.
Like you, I think this is a good thing! Not everyone does.
One more point. I agree that there are a lot of parallels with AI alignment and human/gene alignment. I think you're actually kind of oversimplifying the case by suggesting these subgoals are somehow instilled whole cloth, rather than as a bunch of even further decomposed sub-goals in an elaborate hierarchy. For example, there's the simple urge to have an orgasm, which leads to the temporarily misaligned behavior of masturbation. So it's even more complicated and indirect than just building a 'mate acquisition' drive or module. You kind of hint at this, but it seems underdeveloped.
Anyway, thanks for the read.
> I think you're actually kind of oversimplifying the case by suggesting these subgoals are somehow instilled whole cloth, rather than as a bunch of even further decomposed sub-goals in an elaborate hierarchy.
Totally agree that your description is the correct one.
Regarding women's rights, I agree that this is probably related. Beyond the timelines, it stands to reason that if you're now able to pursue more facets of a good life as opposed to just Parenting, then you would invest somewhat less into Parenting! But is that cultural *evolution*? Maybe this is a boring semantic debate, but usually I think of cultural evolution as people adopting behaviors that ultimately promote reproductive success. If women's rights does the opposite, isn't it better just thought of as falling in the general category of "memetic evolution" or "something we've decided we like"?
(Note for people taking this out of context: Doing things that don't promote reproductive success is fine!)
I would also be very interested in taking lessons from how we (try to) align organizations to human values. I think those are another, great example of a real alignment test case of intelligent systems.
Like: Capitalism wants giant corporations to maximize shareholder value, but instead giant corporations devolve into promotion-packet infighting? Or: Donors want charity to be ruthlessly focused on core mission, but instead charity's mission expands endlessly? I love the idea, but I'm not sure exactly how to identify the analogy. (What exactly is the thing that is being optimized / loss function / optimization algorithm?)
These are also very interesting! But I was thinking of a more direct analogy: organizations are intelligent entities in the world which we try, broadly, in various ways, to align with human values—but the ruthless optimization for profit is rather like a paperclip maximizer, which some would say is actively destroying the planet.
Ha, it's turtles all the way down. I think your description has some truth. But I also think there's some truth in the idea that large companies are in general kind of bad at maximizing shareholder value, simply because it's so hard to align the incentives of large groups of people...
Yeah that’s a good observation, and I don’t think organizations are super intelligences; and probably for this very reason! But then there’s also that time Nestle killed over a million babies, which I think can fairly be described as alignment failure.