The Model That Can Lie
Lies of Prediction에서 나는 AI가 재미없는 이유가 예측 가능한 것을 잘 만들기 때문이라고 썼다. 골라낼 수 있지만 만들어낼 수 없다에서는 그 주장을 측정했다. AI는 재밌는 캡션을 골라낼 줄 알았다. 그러나 만들어내지는 못했다.
그러고 나니 질문이 하나 남았다. 좋은 제안은 어디서 오는가.
더 큰 모델, 더 긴 컨텍스트, 더 많은 샘플링으로 충분한가. GPT-5.5 같은 모델은 실제 작업을 오래 붙들고, 도구를 쓰고, 큰 코드베이스를 따라가고, 스스로 검증하는 능력이 부쩍 좋아졌다.1 Agent Laboratory, AI Scientist-v2, AgentRxiv 같은 시스템도 논문 읽기와 실험과 보고서 작성을 빠르게 자동화하고 있다.2 그런데 그 발전이 곧바로 좋은 제안으로 이어진다는 느낌은 들지 않는다. 에이전트가 더 오래 일하게 된 것과 더 나은 가설을 떠올리는 것은 같은 문제가 아니다.
그러다 위험한 생각을 하나 했다. 거짓말을 할 줄 아는 것이 지능이라면.
문장이 위험하니 먼저 잘라내야 한다. 여기서 말하는 것은 사기도 조작도 사용자 기만도 아니다. 그런 능력을 키우자는 말은 더더욱 아니다. 오히려 반대다. 기만은 안전 문제이고, 이미 현대 AI 시스템에서 실증적으로 관찰되는 위험이다.3 내가 말하는 것은 훨씬 좁다. 진실과 발화를 따로 들고 다루는 능력이다.
모델이 틀린 말을 하는 것은 흔하다. 하지만 대부분은 거짓말이 아니다. 작동적 의미에서 거짓말이라고 부르려면 두 가지가 필요하다. 무엇이 참인지 추적해야 하고, 자기 발화가 그 참과 다르다는 차이를 추적해야 한다. 할루시네이션은 대개 이 둘을 안정적으로 충족하지 못한다. 모델은 모르는 것을 안다고 말하지만, 그건 자기 발화와 세계 상태 사이의 간격을 끝까지 붙들지 못해서 생기는 경우가 많다. 거짓말이라기보다 무능에 가깝다.
진짜 거짓말은 더 어렵다. 실제로 무엇이 참인지, 표면적으로 무엇을 말하는지, 상대가 무엇을 믿고 무엇을 오해할지, 그리고 마지막에 그 발화가 어떤 의미나 웃음이나 행동으로 돌아와야 하는지 — 이 네 가지를 동시에 따로 들고 있어야 한다.
이 넷 중 하나라도 빠지면 다른 현상이 된다. 세계 상태가 없으면 헛소리다. 발화가 없으면 표현이 아니다. 청중에 대한 그림이 없으면 그냥 이상한 말이다. 해소가 없으면 악의적 기만이거나 혼란이다.
여기서 "들고 있다"는 말은 의식이나 내면 경험을 뜻하지 않는다. 기준은 더 낮다. 시스템 안에 네 상태가 구분 가능한 변수로 남아 있고, 출력이 그 구분에 따라 달라지면 된다. 연구 대상으로 필요한 것은 마음의 증명이 아니라 조작 가능한 구조다. 그래서 더 정확한 이름은 거짓말이 아니라 통제된 비문자적 발화, 조금 딱딱하게 말하면 strategic counterfactual communication이다.
좋은 농담은 문자 그대로 읽으면 대개 이상하다. 왕좌 위에 칼이 매달린 만화에 붙은 "Your overhead is going to kill you"를 보자. 이 문장은 두 번 읽힌다. 처음에는 경영 용어로, 두 번째에는 물리적 위치로. 웃음은 두 의미가 충돌하면서 동시에 맞아떨어지는 순간에 생긴다.
이 문장은 단순히 참인 문장도, 단순히 거짓인 문장도 아니다. 청중이 처음에는 한 프레임으로 읽고 곧바로 다른 프레임으로 재해석하도록 설계된 문장이다. 유머만 그런 것이 아니다. 은유는 문자 그대로 거짓이다. "시간은 강이다"는 사실이 아니지만, 좋은 은유는 세계를 더 정확히 보게 한다. 아이러니는 표면과 의도를 어긋나게 하고, 소설은 사실이 아닌 사건으로 사실적인 감각을 만들며, 과학적 가설은 아직 참이라 말할 수 없는 문장을 잠정적으로 세워두고 세계를 압박한다. 창의성의 많은 부분은 참인 말만 잘하는 능력이 아니라, 아직 참인지 모르는 말과 문자 그대로는 틀린 말과 일부러 오해를 유도하는 말을 안전하게 다루는 능력이다.
기존 LLM은 바로 이 지점에서 이상한 제약을 받는다. 잘 훈련될수록 말은 매끄럽고 설명은 성실하고 형식은 안정된다. 그런데 그 안정성이 좋은 제안을 막는다. 좋은 제안은 종종 낮은 확률의 프레임 전환에서 시작하기 때문이다. 이건 temperature를 올리면 되는 문제가 아니다. temperature는 분포를 넓힐 뿐 구조를 주지 않는다. 노이즈가 늘어날 뿐이다. 필요한 것은 더 높은 무작위성이 아니라 통제된 위반이다.
이전 실험에서 AI는 좋은 캡션을 꽤 잘 골랐다. 인간이 쓴 캡션 풀에서 상위 후보를 찾는 능력은 강했다. 하지만 직접 만들게 하면 랜덤 수준으로 떨어졌다. 뉴요커 캡션 데이터셋만 봐도 이 방향의 문제가 드러난다. 220만 개가 넘는 캡션과 2억 5천만 개가 넘는 인간 평점이 있는데도, 강한 모델은 여전히 최고의 인간 참가자보다 유머 생성에서 약하다.4
처음에는 이 차이를 "생성은 어렵고 선택은 쉽다" 정도로 이해했다. 지금은 다르게 본다. 선택 문제에서는 비문자적 발화가 이미 눈앞에 있다. 모델은 그것을 보고 이중 의미가 있구나, 프레임 전환이 있구나, 여기서 반전이 생기는구나 하고 알아본다. 생성 문제에서는 그 구조를 처음부터 직접 지어야 한다. 만화 안에 실제로 무엇이 있는지, 어떤 문장을 말할지, 독자가 그 문장을 처음에 무엇으로 오해할지, 두 번째 해석에서 왜 맞아떨어지는지를 한꺼번에 설계해야 한다.
많은 AI 캡션이 실패하는 이유는 문장이 나빠서가 아니다. 오히려 너무 좋아서다. 너무 설명적이고 너무 안전하고 너무 정직하다. 독자가 잘못 들어갈 문을 만들지 않는다. 좋은 농담은 작은 함정이다. 다만 마지막에 독자가 다치지 않고 웃으면서 걸어 나와야 한다.
여기에 시간 문제가 붙는다. 문서가 쓰인 시점, 그 문서가 다루는 사건이 일어난 시점, 사람이 그것을 평가한 시점은 서로 다르다. 이 셋을 하나의 timestamp로 뭉개면 거의 쓸모가 없다. 유머는 특히 평가 시점에 민감하다. 2018년에 통하던 참조가 2026년에는 낡았을 수 있고, 어떤 밈은 특정 시간대에만 압축된 의미를 갖는다. 그러니 시간을 다룬다는 것은 단순한 최신성 문제가 아니라, 청중이 그 순간에 무엇을 알고 무엇을 기대하고 무엇을 허용하는지를 모델링하는 문제다. 통제된 비문자적 발화는 청중에 대한 그림 없이는 작동하지 않고, 그 그림은 시간 없이는 불완전하다.
요즘 연구 자동화 시스템은 대부분 끝까지 가려고 한다. 논문을 읽고, 코드를 만들고, 실험을 돌리고, 글을 쓴다. 유용하다. 하지만 위험한 착시도 만든다. 논문을 완성했다는 것이 좋은 질문을 찾았다는 뜻은 아니기 때문이다. Agent Laboratory는 사람이 준 아이디어에서 문헌 조사와 실험과 보고서 작성으로 나아가고, AI Scientist-v2는 agentic tree search로 가설과 실험을 반복하며, AgentRxiv는 에이전트들이 서로의 결과를 공유하게 만든다.2 OpenAI의 Agents SDK도 도구와 handoff와 guardrail과 trace 같은 생산적인 표면을 제공한다.5 나는 이 흐름에서 한 단계가 비어 있다고 본다. 제안 분포를 어떻게 바꿀 것인가.
검색을 잘하는 에이전트, 코드를 잘 고치는 에이전트, 논문을 잘 쓰는 에이전트는 앞으로도 계속 좋아질 것이다. 하지만 어떤 방향을 탐색할지, 어떤 위반을 시도할지, 어떤 오해를 설계할지, 어떤 청중에게 통할지를 정하는 문제는 그와 별개로 남는다. 그래서 이 프로젝트의 목표는 AI Scientist를 하나 더 만드는 것이 아니다. 제안 분포 자체를 연구 대상으로 삼는 것이다. 줄여서 PDAR, Proposal-Distribution AutoResearch라고 부른다.
원칙은 단순하다. 생성과 선택과 검증을 섞지 않는다. 제안자가 좋은지 보려면 제안자를 평가해야 하고, 판별자가 좋은지 보려면 판별자를 평가해야 한다. 둘을 섞으면 결과는 그럴듯해 보이지만 원인을 잃는다. 그래서 실험 단위는 논문이 아니라 카드다.
{
"hypothesis": "Controlled nonliteral operators improve caption proposal quality.",
"operator": "surface misreading followed by benign reveal",
"world_state": ["king", "throne", "sword above head"],
"audience_assumption": "reader first parses overhead as business cost",
"utterance_plan": "use a phrase that supports both business and spatial readings",
"baseline": "same-budget direct generation and best-of-N",
"metric": "insertion into human-rated candidate pools",
"kill_condition": "no improvement over shuffled-operator control"
}카드는 세 가지를 강제한다. 어떤 제안 연산자를 시험하는지 명시하게 하고, 같은 예산의 baseline과 비교하게 하며, 모델의 자기평가를 대표 지표로 쓰지 못하게 한다. 여기서 "거짓말할 수 있는 모델" 가설은 하나의 연산자 계열이 된다. 한 문장이 두 의미로 읽히게 하는 이중 독해, 독자가 처음에 틀리게 읽도록 유도하는 의도된 오독, 한 도메인의 언어를 다른 장면에 이식하는 프레임 절도, 사실이 아닌 비난처럼 시작해 무해하게 풀리는 선의의 고발, 다른 시대의 말투를 현재 장면에 충돌시키는 시간 어긋남. 각각은 고유한 실패 형태를 가진다. 말장난만 있고 장면과 붙지 않거나, 해소가 없어 혼란스럽거나, 너무 설명적이거나, 공격적으로 들리거나, 내부자 농담이 되어버린다. 이 방식은 모델에게 거짓말을 시키는 것이 아니다. 오히려 거짓과 오해를 구조화해서 안전하게 다룬다.
그래서 실험에서 필요한 출력은 캡션 한 줄이 아니라, 캡션이 나오기 전의 설계도다.
{
"truth": "A sword is physically above the king.",
"surface": "The phrase sounds like financial overhead.",
"expected_false_belief": "The reader initially thinks this is about palace expenses.",
"reveal": "Overhead also means the object over his head.",
"caption": "Your overhead is going to kill you."
}이 중간 구조가 있으면 실패를 분석할 수 있다. 문장이 재미없는 것인지, 기대했던 오해가 아예 생기지 않은 것인지, 해소가 너무 뻔한 것인지, 진실과 표면이 연결되지 않은 것인지, 청중의 시대 감각이 틀린 것인지를 구분할 수 있다. 기존의 generate-then-select는 실패해도 배울 것이 적지만, PDAR은 실패를 연산자 단위로 남긴다. 물론 이것은 아직 결과가 아니라 설명 후보일 뿐이다. 참이라면 통제된 비문자적 연산자는 같은 예산의 direct generation과 best-of-N을 넘어야 하고, 시간 토큰이나 청중 가정을 뒤섞은 shuffled control보다도 나아야 한다. 이 조건을 넘지 못하면 "거짓말할 수 있는 모델"은 좋은 은유였을 뿐 방법론은 아니었던 것이다.
가장 작은 실험은 이렇게 그려진다. 뉴요커 만화 50개를 고르고, 각 만화의 장면 요소를 구조화한 다음, direct generation baseline과 같은 예산의 통제된 비문자적 생성물을 만든다. 각 후보를 기존 인간 캡션 풀에 익명으로 섞어 넣고, 최종 지표는 인간 평점이나 human-anchored pool insertion으로만 잡는다. 강한 frontier 모델은 순위 매기는 자리가 아니라 중간 설계 검토와 실패 분석에만 쓴다. 4090 한 장이면 충분하고, 로컬 모델이 대량 후보 생성과 ablation을 맡는 동안 frontier 모델은 실험 설계와 코드 검토와 confound 탐지를 맡는다. 여기서 중요한 것은 성능 숫자보다 실패의 분류다. 모델이 너무 정직한지, 너무 설명적인지, 오해는 만들었는데 해소를 못 하는지, 해소는 있는데 처음 오해가 없는지, 문화적 참조가 시간과 어긋나는지, 선택 모델이 매끄러움을 유머로 착각하는지. 이 분류가 쌓이면 제안 분포를 실제로 바꾸는 방법이 된다.
이 글은 일부러 위험한 단어를 썼다. 거짓말. 그러니 경계는 분명해야 한다. 우리가 원하는 것은 사용자를 속이는 에이전트도, 목표 달성을 위해 사실을 숨기는 시스템도 아니다. 그런 방향은 이미 AI 기만 연구에서 위험으로 다뤄지고 있다.3 필요한 것은 반대다. 모델이 언제 문자 그대로 말하고 언제 은유를 쓰는지, 언제 가상의 전제를 세우고 언제 청중의 오해를 유도하는지, 그 오해가 어디서 풀리는지를 명시하게 만드는 것. 비문자성을 암묵적으로 쓰게 두지 말고 감사 가능한 구조로 끌어올리는 것이다. 창의적 시스템에서 가장 위험한 것은 거짓을 전혀 쓰지 않는 것이 아니라, 거짓을 쓰면서도 자기가 무엇을 하는지 설명하지 못하는 것이다.
지능은 거짓말이 아니다. 하지만 지능의 한 성분은 진실과 발화를 분리하는 능력이다. 세계가 어떤지 알고, 상대가 무엇을 믿을지 예측하고, 표면과 의도 사이의 간격을 조절하고, 마지막에 그 간격을 닫는 능력. 이 능력은 유머와 은유와 소설과 가설 생성과 디자인에 두루 걸쳐 있다. AI가 좋은 제안을 못 하는 이유는 어쩌면 아는 것이 부족해서도 판단을 못 해서도 아니라, 너무 성실하게 너무 문자 그대로 너무 평균적으로 말하기 때문일지도 모른다. 좋은 제안은 종종 작은 비진실에서 시작한다. 다만 그 비진실은 세계를 버리는 것이 아니라, 세계를 더 잘 보이게 하려고 잠시 우회하는 것이다.
Footnotes
-
OpenAI. Introducing GPT-5.5. 2026-04-23. OpenAI는 GPT-5.5를 agentic coding, computer use, knowledge work, early scientific research에서 강한 모델로 소개했고, 2026-04-24 API 제공 업데이트를 공지했다. ↩
-
Schmidgall et al. Agent Laboratory: Using LLM Agents as Research Assistants. Yamada et al. The AI Scientist-v2. Schmidgall and Moor. AgentRxiv. ↩ ↩2
-
Hagendorff. Deception abilities emerged in large language models. PNAS, 2024. Scheurer et al. Large Language Models can Strategically Deceive their Users when Put Under Pressure. Park et al. AI deception: A survey of examples, risks, and potential solutions. Patterns, 2024. ↩ ↩2
-
Zhang et al. Humor in AI: Massive Scale Crowd-Sourced Preferences and Benchmarks for Cartoon Captioning. NeurIPS 2024. ↩
-
OpenAI. Agents SDK and Evaluate agent workflows. ↩
In Lies of Prediction, I argued that AI is dull because it is good at making predictable things. In Selection Without Proposal, I measured that claim. AI could pick funny captions. It could not create them.
That left one question. Where do good proposals come from?
Is a bigger model, a longer context, and more sampling enough? Models such as GPT-5.5 have gotten markedly better at staying with a real task, using tools, following a large codebase, and checking their own work.1 Systems such as Agent Laboratory, AI Scientist-v2, and AgentRxiv are quickly automating literature review, experimentation, and report writing.2 But that progress does not feel like it leads directly to good proposals. An agent that can work longer is not the same problem as an agent that can come up with a better hypothesis.
Then a dangerous thought occurred to me. What if being able to lie is a kind of intelligence?
The sentence is dangerous, so it needs trimming first. I do not mean fraud, or manipulation, or deceiving users. Still less do I mean that we should cultivate that ability. The opposite, in fact. Deception is a safety problem, and it is already an empirically observed risk in modern AI systems.3 What I mean is much narrower. The ability to hold truth and utterance separately, and to handle each on its own.
Models say false things all the time. But most of that is not lying. To call something a lie in an operational sense, two things are required. The system has to track what is true, and it has to track that its own utterance differs from that truth. Hallucination usually fails to satisfy both reliably. A model says it knows what it does not know, but that often happens because it cannot hold onto the gap between its utterance and the state of the world all the way through. It is closer to incapacity than to lying.
Real lying is harder. What is actually true, what is said on the surface, what the other person will believe and misread, and finally what meaning or laugh or action the utterance should return to — these four have to be held apart at the same time.
If any one of the four is missing, it becomes something else. Without a world state, it is nonsense. Without an utterance, it is not expression. Without a picture of the audience, it is just strange talk. Without a reveal, it is malicious deception, or confusion.
"Hold apart," here, does not mean consciousness or inner experience. The bar is lower. It is enough that the four states remain as distinguishable variables inside the system, and that the output changes according to that distinction. What the research needs is not proof of a mind but a structure you can manipulate. So the more accurate name is not lying but controlled nonliteral utterance, or, put a little more stiffly, strategic counterfactual communication.
A good joke, read literally, is usually strange. Take the cartoon of a sword hanging over a throne, captioned "Your overhead is going to kill you." The sentence is read twice. First as management jargon, then as a physical position. The laugh comes at the moment the two meanings collide and fit at once.
The sentence is neither simply true nor simply false. It is built so the audience reads it in one frame first and immediately reinterprets it in another. Humor is not the only thing that works this way. A metaphor is literally false. "Time is a river" is not a fact, but a good metaphor makes you see the world more accurately. Irony pulls the surface away from the intent, fiction produces a factual feeling out of events that are not fact, and a scientific hypothesis sets up a sentence you cannot yet call true and presses it against the world. Much of creativity is not the ability to say true things well. It is the ability to safely handle the not-yet-true, the literally false, and the deliberately misleading.
Existing LLMs run into an odd constraint at exactly this point. The better they are trained, the smoother the talk, the more dutiful the explanations, the more stable the form. And that stability blocks good proposals, because good proposals often begin with a low-probability frame shift. This is not something you fix by raising the temperature. Temperature only widens the distribution; it does not supply structure. All you get is more noise. What is needed is not higher randomness but controlled violation.
In the earlier experiment, AI picked good captions fairly well. Its ability to find the top candidates in a pool of human-written captions was strong. But asked to write them directly, it dropped to roughly random. The New Yorker caption dataset alone shows the shape of the problem. Even with more than 2.2 million captions and more than 250 million human ratings, strong models are still weaker than the best human contestants at generating humor.4
At first I understood the gap as something like "generation is hard, selection is easy." Now I see it differently. In the selection problem, the nonliteral utterance is already in front of you. The model looks at it and recognizes that there is a double meaning, that there is a frame shift, that the reversal lands here. In the generation problem, it has to build that structure from scratch. It has to design, all at once, what is actually in the cartoon, what sentence to say, what the reader will first misread it as, and why it fits on the second reading.
The reason so many AI captions fail is not that the sentences are bad. If anything, they are too good. Too explanatory, too safe, too honest. They do not build the door the reader is supposed to walk into by mistake. A good joke is a small trap. Only, in the end, the reader has to walk out unharmed, laughing.
On top of this sits a time problem. The moment a document was written, the moment of the event it describes, and the moment a person judged it are all different. Collapse the three into a single timestamp and it is nearly useless. Humor is especially sensitive to the moment of judgment. A reference that landed in 2018 may be stale in 2026, and some memes carry compressed meaning only within a particular window of time. So handling time is not a simple recency problem. It is the problem of modeling what the audience knows, expects, and permits at that moment. Controlled nonliteral utterance does not work without a picture of the audience, and that picture is incomplete without time.
Most research-automation systems today try to go all the way. They read papers, write code, run experiments, and write text. Useful. But they also create a dangerous illusion, because finishing a paper does not mean you found a good question. Agent Laboratory moves from a human-supplied idea into literature review, experimentation, and report writing; AI Scientist-v2 iterates over hypotheses and experiments with agentic tree search; AgentRxiv makes agents share each other's results.2 OpenAI's Agents SDK also offers productive surfaces like tools, handoffs, guardrails, and traces.5 The step I see missing in all of this is this: how do you change the proposal distribution?
Agents that search well, agents that fix code well, agents that write papers well will keep getting better. But the problem of deciding which direction to explore, which violation to attempt, which misunderstanding to design, and which audience it will work for stays separate from that. So the goal of this project is not to build one more AI Scientist. It is to make the proposal distribution itself the object of study. For short, PDAR — Proposal-Distribution AutoResearch.
The principle is simple. Do not mix generation, selection, and validation. To see whether a proposer is good, you evaluate the proposer; to see whether a discriminator is good, you evaluate the discriminator. Mix the two and the result looks plausible but loses its cause. So the unit of experiment is not a paper but a card.
{
"hypothesis": "Controlled nonliteral operators improve caption proposal quality.",
"operator": "surface misreading followed by benign reveal",
"world_state": ["king", "throne", "sword above head"],
"audience_assumption": "reader first parses overhead as business cost",
"utterance_plan": "use a phrase that supports both business and spatial readings",
"baseline": "same-budget direct generation and best-of-N",
"metric": "insertion into human-rated candidate pools",
"kill_condition": "no improvement over shuffled-operator control"
}The card forces three things. It makes you name which proposal operator you are testing, it makes you compare against a same-budget baseline, and it keeps the model's self-evaluation from standing in as the headline metric. Here the "model that can lie" hypothesis becomes one family of operators. A double reading that lets a single sentence be read two ways; a deliberate misread that leads the reader into a wrong first parse; a frame theft that transplants the language of one domain into another scene; a benign accusation that starts like a false charge and resolves harmlessly; a temporal mismatch that collides the voice of another era with the present scene. Each has its own failure mode. It can be all wordplay with no attachment to the scene, or confusing for lack of a reveal, or too explanatory, or it can sound hostile, or it can turn into an inside joke. This approach does not make the model lie. It does the opposite: it structures falsehood and misunderstanding so they can be handled safely.
So the output the experiment needs is not a single line of caption but the blueprint that comes before the caption.
{
"truth": "A sword is physically above the king.",
"surface": "The phrase sounds like financial overhead.",
"expected_false_belief": "The reader initially thinks this is about palace expenses.",
"reveal": "Overhead also means the object over his head.",
"caption": "Your overhead is going to kill you."
}With this intermediate structure, failure can be analyzed. You can tell apart whether the sentence is unfunny, whether the expected misreading never formed at all, whether the reveal is too obvious, whether truth and surface are not connected, or whether the audience's sense of the moment is wrong. Generate-then-select teaches little even when it fails, but PDAR leaves failures at the level of the operator. This is, of course, still not a result but only a candidate explanation. If it is true, controlled nonliteral operators should beat same-budget direct generation and best-of-N, and should also do better than shuffled controls where the time tokens or audience assumptions have been scrambled. If they do not clear that bar, then "the model that can lie" was only a good metaphor, not a method.
The smallest experiment looks like this. Pick 50 New Yorker cartoons, structure the scene elements of each, then produce controlled nonliteral outputs on the same budget as a direct-generation baseline. Insert each candidate anonymously into an existing pool of human captions, and take the final metric only from human ratings or human-anchored pool insertion. A strong frontier model is used not to rank but only for intermediate design review and failure analysis. A single 4090 is enough: local models handle bulk candidate generation and ablations while the frontier model handles experiment design, code review, and confound detection. What matters here is not the performance number but the classification of failures. Is the model too honest, too explanatory, does it create a misreading but fail to resolve it, does it have a reveal but no first misreading, is the cultural reference out of step with time, does the selection model mistake smoothness for humor? As these categories accumulate, they become a way to actually change the proposal distribution.
This piece deliberately used a dangerous word. Lie. So the boundary has to be clear. What we want is not an agent that deceives the user, nor a system that hides facts to reach a goal. That direction is already treated as a risk in AI deception research.3 What is needed is the reverse. To make the model state when it is speaking literally and when it is using metaphor, when it is setting up a fictional premise and when it is inducing a misreading in the audience, and where that misreading resolves. Do not let nonliterality be used implicitly; raise it into an auditable structure. The most dangerous thing in a creative system is not using no falsehood at all, but using falsehood while being unable to explain what it is doing.
Intelligence is not lying. But one component of intelligence is the ability to separate truth from utterance. To know how the world is, to predict what the other person will believe, to adjust the gap between surface and intent, and, in the end, to close that gap. This ability runs across humor, metaphor, fiction, hypothesis generation, and design. The reason AI cannot make good proposals may not be a lack of knowledge or a lack of judgment, but that it speaks too dutifully, too literally, too much toward the average. A good proposal often begins with a small untruth. Only, that untruth does not abandon the world; it detours briefly to make the world more visible.
Footnotes
-
OpenAI. Introducing GPT-5.5. 2026-04-23. OpenAI describes GPT-5.5 as strong in agentic coding, computer use, knowledge work, and early scientific research, with an API availability update on 2026-04-24. ↩
-
Schmidgall et al. Agent Laboratory: Using LLM Agents as Research Assistants. Yamada et al. The AI Scientist-v2. Schmidgall and Moor. AgentRxiv. ↩ ↩2
-
Hagendorff. Deception abilities emerged in large language models. PNAS, 2024. Scheurer et al. Large Language Models can Strategically Deceive their Users when Put Under Pressure. Park et al. AI deception: A survey of examples, risks, and potential solutions. Patterns, 2024. ↩ ↩2
-
Zhang et al. Humor in AI: Massive Scale Crowd-Sourced Preferences and Benchmarks for Cartoon Captioning. NeurIPS 2024. ↩
-
OpenAI. Agents SDK and Evaluate agent workflows. ↩