LLM benchmark progress, the "easy" problem, and the persistence of the "hard" problem
The "Hard Problem" hasn't gotten any easier for LLMs in the last 12 months.

I truly feel that flagship model LLMs’ comprehension and general everyday competencies have become less capable. Ethan Mollick1 and others show many wonderful things that are fascinatingly capable, even as one-shot prompts. Yet I find most of the important work I'm doing and advising others to do is not of the one-shot nature; it is slow, iterative, and requires much exploration. But, this even holds to some incredibly basic auto-suggested changes and corrections.
I find myself2 continuously being in dismay when corrections that were previously streamlined are now clunky and require more revising. I understand several performance benchmarks show substantial improvement, but I don't have much to say in favor of the general use for most LLMs as becoming more cognitively fluent or more adapted to the user (or to 'me', in particular.)
When I structure it to do what it should do at a high level, it works well; if you provide the structure or give it a blank canvas with defined constraints; or, if the standard is merely that something interesting comes out and the specifics are negotiable, these outputs can be genuinely striking.
What I haven't seen improve is the underlying robustness; nothing I've experienced is particularly improved in what I'd say is actual cognitive functioning or becoming less-brittle in its contextual awareness. It is still, as it was listed some time ago, "an intern that needs much guidance on how to do things" - that intern is now a fantastic coder, and can pull sophisticated-sounding information and citations; there is some element of brilliance there.
Yet for the kind of work that I find most critical and most important and requires the most actual consideration, or deep work, I am left feeling its almost taken a step back in the last 12 months, rather than improved.3
I know this is probably a controversial hot take-like post, or something you'd hear from a card-carrying member of the AI Hater Club; that's for you to determine.
What I'd add is that I would prefer it weren't true.
Hand-holding a model through a critical conceptualization is time I would rather spend elsewhere, and I keep being reminded that the limits of the intellectual partnership are what they are.
Easy and Hard problems for in AI for Science
My reference of the day: This piece by Ruairidh Battleday and Sam Gershman gets at some of it - the difference between the "hard problem" (not consciousness!) and "easy problem."
In other words, if you're research or business aims are "easy problem" related - this is fantastic times for many AI related implementations, for sure.
But, indeed, "Solving the hard problem is beyond the capacities of current algorithms for scientific discovery because it requires continual conceptual revision based on poorly defined constraints"; even though this paper was made perhaps over a year ago and published earlier this year proper, I would say this still holds.
PS: We had the opportunity to cover this in the Cognition Futures (JOPRO + Orthogonal Research and Education Laboratory) Reading Group recently, which will be posted on our YouTube soon.
Note: This post is adapted from my original LinkedIn post
Someone I widely recommend and appreciate the sharing of his views and continuous experiments in those spaces.
For disclosure, I’ve used the latest models from OpenAI, Anthropic, and Google relatively consistently over thee last 18 months. I do this in a professional sense in terms of consulting, technology & executive strategy, as well as in a more educational or research-specific sense as well.
My one caveated consideration here: if the models are getting better at large, perhaps there is an intentional gimping of this for general consumer-grade uses, but higher-end or more selective models are hoarding the capacity that I’m pointing out as lacking.



