A few days ago, a shared ChatGPT conversation involving Terence Tao — one of the most decorated mathematicians alive — circulated on Hacker News. The thread showed an AI system producing a confident, elaborately reasoned claim about the Jacobian Conjecture, a famously unsolved problem in mathematics. The claim was wrong. The lesson isn't that AI is useless. It's that AI can construct plausible-sounding justification for incorrect conclusions in ways that are very hard to detect unless you already know the answer.
If a man who won the Fields Medal needs to carefully interrogate an AI's mathematical output, your household is not immune to the same failure mode.
What's actually changing
The useful thing about this episode isn't the math. It's the structure of the error.
Modern large language models don't reason the way a careful person reasons. They generate text that is statistically coherent with their training data. When the output happens to align with truth, that's partly because correct things are frequently written down, not because the model is checking its work. On a well-worn topic — standard recipes, common tax rules, medication dosage guidelines — that approach works often enough to feel reliable. On anything unusual, novel, contested, or at the edge of a domain, the model can confidently produce fiction that reads exactly like fact.
The Hacker News discussion around the Tao exchange made this visible in a high-stakes, high-credibility context. But families are running the same risk every time they use AI output to make decisions without a second checkpoint: what to do when a prescription conflicts with another medication, whether a lease clause is legal in their state, how to safely handle a household emergency, which evacuation route to take.
The capability of these tools has grown faster than public understanding of where they fail. That gap is the actual preparedness problem.
What we'd actually do
Categorize your AI use by consequence, not convenience. Start by writing down the last ten things you used an AI tool to answer. Sort them into two columns: low-stakes (recipe substitutions, gift ideas, draft emails) and high-stakes (medical, legal, financial, safety). Low-stakes use needs no new process. High-stakes use does. This single habit change — pausing to ask "what happens if this is wrong?" — is the whole ball game.
For anything in the high-stakes column, require a named, verifiable second source. Not "the AI confirmed it" or "I asked again and got the same answer." A second AI query is not a second source — it's the same model saying the same thing twice. You want a human professional, an official agency page, a dated publication from a known institution. This costs time, not money.
Build a short household reference list of domains where AI has failed you or people you trust. Keep it somewhere findable — a note on your phone, a sticky inside a cabinet. Add to it when you see a credible public error like the Tao example. The list isn't a reason to stop using AI. It's a calibration tool. When your list grows heavy in medical and legal domains, that tells you where to slow down.
Talk to your kids about this explicitly, by name. Teenagers are using AI for homework, health questions, and social decisions. The concept they need isn't "AI is bad" — it's "confident-sounding is not the same as accurate." The Tao story is a good teaching case because it's concrete: even world-class experts have to check the machine's work.
Keep critical decision documents in a format you can read without an AI intermediary. Emergency contacts, insurance policy numbers, local shelter locations, medication schedules — these should be on paper or in a simple text file you control. If the tool you rely on to synthesize information is the thing that fails, you need a fallback that doesn't require another tool.
The bigger picture
AI tools will keep getting better at sounding correct. That's not speculation — it's what the last several capability jumps have looked like. But sounding correct and being correct are different things, and the gap between them doesn't close on a predictable schedule. Preparedness, at the household level, has always been about building systems that don't collapse when one component fails. AI is now a component in many households' decision-making. Treat it like any other single point of failure: useful, worth having, not worth depending on entirely.
The goal isn't to distrust everything. It's to build the habit of knowing which decisions need a second opinion — and having a plan for getting one.





