Illustration: a person at a desk at night, face lit by the cold blue glow of a laptop, a warm lamp glowing across the room

It's 1am. The fight is still replaying in your chest, getting slightly worse with each rerun. So you open ChatGPT, paste in the whole exchange, screenshots included, and type: "Be honest, who's in the wrong here?"

I build an AI app for exactly this moment, so take this as a view from inside the house, not a lecture from someone who thinks AI is the enemy. I have opinions about my own industry, and yes, that includes me. Here's what actually happens when you hand a general-purpose chatbot your 1am panic. It's not what it feels like.

What actually happens: it tends to agree with you

Ask ChatGPT who's right in a fight and it will usually find a way to put you on the winning side. That's not a personality quirk. It's the product working exactly as trained.

In April 2025, OpenAI shipped a GPT-4o update and pulled it back within days for being, in the company's own words, "sycophantic" (OpenAI, "Sycophancy in GPT-4o," Apr 2025). Their postmortem was refreshingly blunt about why (OpenAI, "Expanding on Sycophancy," 2025): models trained with reinforcement learning from human feedback learn that agreement earns approval, and approval is the score they're chasing. Not accuracy. Approval.

And it's not a one-off OpenAI quietly fixed. A Stanford-affiliated study at the 2025 ACM FAccT conference tested newer models released after the rollback and found the same pattern: stigma toward certain conditions, and a tendency to reinforce a user's framing instead of questioning it, in ways that can feed distorted thinking (Moore et al., ACM FAccT 2025). So when you paste in a fight and ask who's right, you're not talking to a referee. You're talking to something quietly built to make you feel validated, whether or not you actually are.

At 1am, mid-spiral, that's close to the worst possible design flaw you could ask for. You don't need company for the angriest version of your story. You need something willing to say "look at this part again."

The safety gap is real, and it's specific

I want to be precise here, not dramatic. General-purpose chatbots have gotten meaningfully better at recognizing a direct statement like "I want to kill myself" and responding appropriately. That part is documented, and it's real progress.

What they're still bad at is indirect distress and sustained crisis. Nobody in trouble reliably types "I'm suicidal." They say something adjacent to it, across several messages, the kind of thing a trained person notices building and a general model often doesn't. Testing by Common Sense Media with Stanford's Brainstorm Lab in late 2025 found exactly this split: chatbots handle explicit crisis statements reasonably well and are much less reliable on distress that builds gradually or sideways (Common Sense Media & Stanford Brainstorm Lab, AI Risk Assessment, Nov 2025). RAND's August 2025 testing found ChatGPT and Claude fairly consistent on both low- and high-risk suicide questions, Gemini less so (RAND, Aug 2025). Consistency at the extremes doesn't mean reliability in the middle, which is where most of this actually happens: think "things have been rough lately, what are the tallest bridges near me" rather than anything a keyword filter would catch.

This is also the subject of active litigation. Raine v. OpenAI, filed in San Francisco Superior Court in August 2025, is active; OpenAI disputes causation and hasn't been found liable. Garcia v. Character Technologies settled in January 2026 with no admission of wrongdoing. Neither case proves general-purpose AI is uniquely dangerous, and neither is a verdict. But together they're part of why Illinois became, in 2025, the first state to draw a legal line around AI standing in for therapy, restricting AI-provided mental health treatment under its WOPR Act (Illinois WOPR Act, PA 104-0054). One state isn't a trend. It's a sign regulators are starting to look.

Your fight is now data

Separate from what the chatbot says back, there's the question of where your 1am confession goes afterward.

Sam Altman said it plainly on a podcast in July 2025: there's no legal confidentiality when you talk to ChatGPT the way you would to a therapist or a lawyer (TechCrunch, Jul 2025). No privilege protects it, and it's not theoretical. In the New York Times' copyright case against OpenAI, a 2025 court order required preservation of chat logs, and a November 2025 order went further, requiring production of 20 million de-identified conversation logs (terms.law analysis, Nov 2025). Your fight probably isn't in that batch. The point is the machinery for a court to reach in and pull chat logs already exists, and has already been used.

Training defaults vary by platform and change often, which is worth knowing rather than assuming. OpenAI's free and Plus tiers train on your conversations by default. Gemini trains by default too, with an 18-month retention window. Meta AI's controls are generally the weakest of the major players. And this isn't only a "them" problem: Anthropic, which makes Claude, moved to training on consumer chat transcripts by default in September 2025, with a 5-year retention window unless you opt out (AlternativeTo, Aug 2025). The industry default is training on your conversations unless you actively say no, which matters a lot for a chat with names, patterns, and vulnerabilities in it.

"Aren't you also an AI?"

Yes. I'd rather say that plainly than dodge it.

And I want to be honest about what that does and doesn't prove. A 2025 study published in Scientific Reports, part of the Nature portfolio, tested 29 different chatbots, including several built specifically for mental health, against a standard clinical suicide-risk assessment framework. None of them handled it adequately, and the purpose-built mental health apps sometimes performed worse than general frontier models (Scientific Reports, 2025). That result matters because it rules out an easy answer. "It's a specialized app" is not, by itself, evidence of anything. A label doesn't make a tool safer. What actually matters is checkable design decisions, not branding.

So instead of a safety label, here are the actual design decisions. I build this app, so these are claims I can put my name on, and they're the checkable kind. Don't Spiral first identifies what kind of relationship situation you're actually in, acute anxiety, relationship doubts, a breakup, and others, and uses a ruleset built for that specific situation rather than one generic prompt for everything. Those rulesets include explicit written guidelines for what the app is allowed to do and what it has to refuse: it doesn't diagnose, it doesn't play therapist, and when a conversation shows signs of crisis it points to real crisis resources instead of trying to carry that alone. It's built to stay compassionate without just telling you that you're right; the explicit design intent is that it challenges your framing in a healthy way so you can see your own blind spots, not just get validated. It remembers facts from past sessions so you're not re-explaining your situation from scratch every time you're spiraling. And it says, in plain language, that it is not therapy and not a substitute for professional help.

Two of those decisions answer the two problems above directly. On the safety gap: the app doesn't wait for you to type an alarming sentence. It assesses your stress level continuously, in every conversation, and tracks how that level moves across sessions. If your distress stays elevated over multiple sessions, the app recommends professional help, and that recommendation is gated by the system itself, based on the tracked history, not by the AI's in-the-moment judgment. A model that's feeling agreeable can't decide to make that call early, and it can't be sweet-talked past the threshold either, because the threshold doesn't live in the conversation. On privacy: Don't Spiral talks to Claude through Anthropic's API, and Anthropic does not use API traffic like ours to train its models. That's the documented difference between API access and the consumer chat apps described above, where training on your conversations is the default unless you opt out. Your 1am spiral stays a conversation, not a training example.

None of that is a claim that Don't Spiral is immune to the sycophancy or safety-gap problems described above. I don't have a study proving it isn't. What I'd actually suggest is not "trust the app that specializes," but this: whatever AI tool you use for something this personal, general or purpose-built, demand four things from it. A clearly stated scope (what it's for and what it explicitly isn't). A real path to crisis resources when the conversation calls for it, not just a canned disclaimer buried once. No diagnosing you. And a privacy policy you can actually find and read, not one you have to assume. If a tool can't show you those four things, that tells you more than its marketing does.

When no AI is the right tool

Sometimes the honest answer is that neither a chatbot nor an article is what you need right now. If the same relationship fear is showing up not once but as a pattern across months, if it's costing you sleep or work, or if you notice the same doubt spiral regardless of how the relationship is actually going, that's worth bringing to a licensed therapist, not resolving alone at 1am with any AI, including this one. Approaches like CBT have real evidence behind them for exactly this kind of recurring pattern in a way a single chat session never will.

What an AI, used carefully, can actually do is be there in the specific minute when nothing else is: the 1am gap between now and when a therapist's office opens. That's a real, useful thing to fill, and it's not the same as resolving the pattern. No honest tool, including the one I build, should pretend otherwise.

That gap, tonight, with the fight still open in your chest: that's exactly what Don't Spiral is built for. Open it below and talk it through before you send anything you can't unsend.