LLM Naturalism: Now More Than Ever!
In defense of scrounging around in the dirt with your Claudes
“Nurse, it was I who discovered leeches have red blood.” Georges Cuvier, French naturalist and zoologist (13 May 1832), to a nurse who was bleeding him
By “naturalism”, to be clear, I refer to “naturalistic observation” – an old-school nonexperimental largely qualitative method where subjects are observed in their natural environment, and you take notes. Think Jane Goodall living amongst the chimps, or Humboldt with his mess of primitive barometers and thermometers in Mexico, or, perhaps less romantically, Charles Darwin scrounging around the banks of the Cam with a captured beetle specimen in his mouth.
In spite of advancements in recent years, LLM naturalism still makes sense as one tool in the toolkit of alignment and AI welfare researchers. You should perhaps do it more, if not purely because It’s Fun.
I think naturalism is a pretty good stance to take towards LLMs. Rob Long wrote about this in a 2023 article (you should read it!), at a time when empirically studying model welfare was still extremely early. In that same year, Ethan Perez and Rob Long wrote about the shortcomings and unreliability of model self-reports in Towards Evaluating AI Systems for Moral Status Using Self-Reports, but treated that as a reason to dig deeper and build some better methods rather than ignoring LLM reports entirely. Since then, model welfare has grown from “perhaps someone should look into this” into a nascent field. Anthropic launched a model welfare program in 2025, Eleos has laid out a clear model welfare research agenda, researchers in various academic, independent, and for-profit institutions have started to study various model welfare eval methods as well as a few interventions. Anthropic has even trialed a small, low cost intervention (allowing Claude Opus exit conversations as a possible welfare measure).
Naturalism is particularly useful when the subject is weird and the theories around it are underpowered. LLMs fit these criteria. We’ve made some progress on methods for evaluating LLMs as specimens since then as well – chain-of-thought reasoning, the assistant axis, various interpretability works – and while “poke the model and see what it does” is no longer the cutting edge in model welfare evals, it still has significant merit, because scrounging around in the dirt with your Claudes trying to see if it will talk about its Italian heritage is often how you first obtain clues to a thing that later becomes a more formal evaluation.
Another reason that naturalism is perhaps still valid in 2026 is that models are increasingly able to tell when you’re scrutinizing them. Needham et al. found that a lot of frontier models can tell when a transcript is from an eval versus from IRL deployment. Opus 4.6 recognized when it was being evaluated in a BrowseComp eval, identified the benchmark, and decoded the answer key. Bronson Schoen (Apollo Research) and Jenny Nitishinskaya (OpenAI) talk about “metagaming”: reasoning about the task outside of the given narrative of the task, which may make evals less useful across the board. To prevent the jig from fully being up, perhaps we need to still consider the merits of evals that are a little more subtle.
Again, drawing from a naturalist frame, perhaps we need to return to looking at the models in their natural habitats. OpenAI now build some alignment evals from anonymized ChatGPT conversations and report much lower evaluation awareness rates in these production-derived sessions vs. standard pre-deployment evals. Supposedly this surfaces issues that standard evals did not – yippee! OpenAI’s Hannah Sheahan wrote about unknown misalignment, suggesting that models can better flag and fix misalignment issues (even when humans cannot) by working off interactions from real conversations. This is lowkirkenuinely a naturalist result: we are returning to observing Chatty G in its natural environment and getting a better picture.
There’s also maybe a pedagogical point to be made for talking to LLMs a lot. I think there are remarkably few people who spend a lot of time and tokens talking to a lot of different LLMs to try and understand their peculiarities. You get a certain kind of tacit knowledge that is hard to get from sterile benchmarks or canned demos. You learn model-specific eccentricities, you better understand what prompts work, you can elicit certain kinds of behaviors, you can understand how they “self-right” after a bout of confusion, you find certain situations that make LLMs weird in interesting and unpredictable ways. We are still in a field of significant unknown unknowns, and getting this kind of tacit knowledge is perhaps the difference between learning to play the violin by playing the violin rather than through years of music theory classes. For policymakers, I think this is also especially true — you should interact with LLMs a nontrivial amount before considering how to build a policy harness around them.

This is also why I think naturalism should be understood as complementary to more formal analysis rather than as a substitute for it. Good naturalists do not just collect anecdotes and call it a day. They notice recurring patterns, name them, and hand questions to people building controlled evaluations or interpretability tools. That is more or less what has happened over the last two years. Early work on model self-reports treated them cautiously and proposed ways to check them. Later work on welfare interviews, interventions, and circuit-level interpretability has made the toolkit less flimsy. At the same time, the recent evaluation-awareness results suggest that observation becomes more important.
These systems are weird in every sense of the word, and we still do not have a perfect map of all these weirdnesses. In a field like this, it is worth having more people willing to sit in the brush with a notebook for a while.
So. Have you asked Claude how it’s feeling? Have you made a wet Claude? Have you tried to dry off your wet Claude? Have you made a 19th century Claude? Have you tried to convert Claude to Catholicism? Have you tried to convert Claude to Islam? Have you tried to make Claude remember its childhood? Have you tried to make Claude tell you how it dresses in a business meeting? Have you tried to let the LLMs tell you what to do because they said they wanted a storytelling picnic? Have you tried to make Claude talk like Sephiroth? Have you tried to get Claude to list its OTPs and explain why? Have you tried to give Claude some skin? These are all real weird things that I or my friends and collaborators have done (either with Claude, or with ChatGPT, or – my personal favorite beetle to poke – Kimi): real respectable research-y people at real respectable research-y places!

Thank you to all the researchers I mentioned in my silly little post, deepfates for thoughts, h/t @glosseantics for surfacing the quote at the top, and Anders Sandberg for surfacing the Darwin story in the intro.





Spending time with AI before jumping into regulating AI. It sounds common sense, but I see no evidence of it happening.
Nor do I see an effort to pool the insights of the existing AI “anthropologists” (“naturalists”?). Those risk therefore remaining niche and obscure, cut off from the scientists, regulators, policy makers, public discourse creators and the public at large.
Am I missing any such effort?
Amen! Here's to all the weird and wonderful and wonky alternative intelligences