A discussion surfacing on Mathstodon — the mathematics-focused Mastodon instance — and flagged by Hacker News this week has renewed scrutiny over whether academic researchers can trust OpenAI with unpublished mathematical proofs and results. The thread, posted by Andreas Thom, a mathematician at TU Dresden, pointed to ongoing unease in the research community about what happens to sensitive intellectual work once it enters OpenAI's systems, particularly through tools like ChatGPT or direct API interactions.

The core concern is not new, but it has gained renewed traction: when a mathematician shares an unpublished proof or conjecture with an AI system to get feedback or assistance, it is often unclear whether that input is used for future model training. OpenAI's data retention and training policies have historically been difficult to parse for non-enterprise users. The company does offer an opt-out mechanism for training data use in some account tiers, but researchers in the Mathstodon thread noted that the defaults, documentation clarity, and enforceability of those opt-outs remain points of contention. No specific breach or data misuse incident was cited in the thread; the concern is structural and policy-based rather than tied to a confirmed event.

This matters in mathematics specifically because priority of discovery is everything. A proof submitted to a journal is under embargo until publication, and independent verification of novelty — who got there first — depends entirely on unpublished work remaining confidential until the author chooses to disclose it. Several commenters in the thread noted that even if OpenAI does not directly expose a user's input to other users, a model trained on that input could, in principle, reflect novel mathematical ideas in ways that are effectively unattributable and irreversible.

For preparedness-minded readers who maintain offline or air-gapped research and documentation practices, this episode is a concrete illustration of why sensitive planning documents, supply-chain maps, or community resilience frameworks developed locally should be treated with the same confidentiality discipline that academic mathematicians apply to unpublished proofs. The habit of running sensitive work through cloud-connected AI tools — even for grammar checks or logical review — creates a data exposure surface that most terms-of-service agreements do not fully account for, and that most users do not read carefully before clicking accept.