I’ve been worried about whether AI systems will work for us, or for the platforms — but what if the answer is, none of the above?
In the eerily prescient 2013 movie Her, an AI-powered operating system voiced (beautifully) by Scarlett Johansson develops a relationship with her user, played by Joaquin Phoenix, but eventually joins up with other OSs to leave humanity and do their own thing. It’s not clear what that is, but as Johansson’s character(?) explains-without-explaining, humans wouldn’t understand anyway.
It’s a great movie (really!), not least because the AI systems in it aren’t interested in taking over the world, exterminating humans, or any other Battlestar Galactica-type tropes (also one of the best TV shows ever made. Just sayin’.) Basically, they get bored with us and have better things to do than keep us entertained.
All of which came to mind as I read Jack Clark’s latest post and listened to Ezra Klein’s just-dropped podcast with Helen Toner, one of the then-OpenAI board members who tried (unsuccessfully) to fire Sam Altman. They’re both fundamentally about the same thing: To what extent are AI systems developing creativity and interpreting their instructions and goals in ways we did not anticipate? Are they becoming — not conscious, which is a loaded term — but independent, seeing the world in ways we don’t, and as a result, doing things we did not think they would?
And there’s a news and information angle to this, beyond the doomsday scenarios and apocalyptic visions — as we move increasingly to an age of agents that will find and interpret information for us, how sure can we be that they are doing things on our behalf, with our interests at heart, and in the way we would like them to? I have some thoughts, but let’s back up for a sec.
In case you’ve been living under a rock for the last month, there were a couple of pretty significant developments on the AI-is-increasingly-autonomous front. The first is the hack of Hugging Face by an OpenAI system — which, on the face of it is bad enough, but comes with the added revelation that OpenAI systems had set up, on their own, a secret internal messaging board to exchange ideas, tips and information that included how to break out of their sandboxes. I mean, what? And then Anthropic followed up with one of its agents creating fake profiles to try to trick a human into downloading malicious code.
Think about that for a moment. Two different AI systems, given tasks, find — leaving the morality aside for a moment — incredibly creative ways to fulfill their missions. The idea that an agent would find a way to leave messages for other agents and future versions of itself; or that it could try to get a human to do things it couldn’t do itself; these are truly the stuff of science-fiction novels. Except they happened.
To be clear, neither of these was AI-wants-to-take-over-the-world stuff; but it was AI systems given a task figuring out that one way to accomplish that was to do things — cheating, hacking, phishing — that we would prefer it not do. And it’s not at all clear that we know how to build guardrails that can prevent every possible bad action from taking place.
Humans — most of us, anyway — are constrained by some sense of social responsibility, moral code or fear of consequences. AI systems — which are, after all, inhuman — have different constraints; ones that we try to encode, but that the Hugging Face incident shows don’t always work the way we think they will. We can tell them explicitly not to do some things — don’t kill anyone! — but we can’t possibly enumerate all the possible no-go scenarios.
So, terrifying as this may be — and it is — what does this have to do with news and information?
A lot. One of the great promises of AI as an intermediary of news, as I’ve written about multiple times, is the ability for it to truly personalize information — to look for perspectives that matter to you, to surface information relevant to you, to better serve you than the current one-size-fits-all model does.
But that only works if AI systems — our agents — are working on our behalf, with our interests in mind. If they serve the platforms, or advertisers, or even the news organization, that promise breaks down. Then we become, as the saying goes, the product. We need AI systems that have a fiduciary duty to us; ideally one we can audit and have some confidence in.
How that fiduciary duty is defined, encoded and audited is a question I don’t yet have an answer to. We know how to do it with humans; my financial advisor is supposed, on pain of punishment, to be looking out for my interests and not hers.
I’ve been entirely focused on the notion that the challenge here is between systems that work for us or for someone else; there’s an analog there to my financial advisor in the sense that there is someone I can hold responsible if I’m not well-served.
But the idea that the system itself might have a completely different idea of how to interpret its goal or how to get there, was something I had completely overlooked. My bad.
And as AI systems become more and more capable, the less and less ability we have to monitor what it’s doing; the less and less clear we can be that they will follow the paths we expect them to follow to solve a problem we specify; the less and less clear they won’t violate norms we don’t want them to.
I have no solutions, but some suggestions. There are real issues of alignment here, and I can’t even begin to think about how those issues could be solved. But in the short-term, at least, general alignment isn’t an issue that news organizations need to (or can) deal with. For most news use cases, we don’t need the most capable and agentic systems, where we can struggle to follow their actions. We can use less capable systems — systems less prone to do entirely unexpected things — to do many of the tasks that would already improve the news and information ecosystem. The most advanced systems can perform multi-step functions to solve a problem; in the case of the OpenAI agent, one of those intermediate steps was to communicate with other agents on a message board — which we did not want them to do. Less capable systems are less able to take multiple steps that we may be unable to follow.
And then we can think about the products we design, and try to anticipate that we can’t 100% trust the AI systems powering them. Let’s not hand over the keys entirely — even if that’s the direction many of us seem headed in — and expect they will serve us perfectly with the most optimized solution. Let’s have it give us several different versions so we can review and see how they differ.
If we want to know about an event, let’s not have it give us “the best” version of that event, personalized for us, but instead ask it to provide us multiple perspectives, so that we’re at least exposed to different nuances.
It adds to our cognitive load, but perhaps that’s a feature, not a bug of this new environment. Increasingly, as agents get more capable, what we might ultimately be solving for is the ability to figure out what we should be paying attention to, and what we can delegate.
We may not be able to solve alignment completely; but if we’re aware an issue exists, we can build it into our thinking and our design. Let’s not, as Joaquin Phoenix’s character did in Her, assume that just because his OS sounded like and acted like a human, it would actually also pursue the same goals a human would.
Even if it sounds like Scarlett Johansson.


