Skip to main content

Ethyx

Log In
Back to blog

September 9, 2026 · 6 min read

By David Crush

Who owns the work you give an AI?

This week’s fight over a math proof is not really about one Millennium Prize problem. It is about what happens when unpublished work lives inside a company’s product—and who gets to improve a model from it.

That question is a large part of why I started Ethyx.

What was reported

NYU mathematician Tristan Buckmaster and Levent Alpöge posted results on problems related to Navier–Stokes after using tools including OpenAI’s Codex and Anthropic’s Claude. Buckmaster says information about their unpublished progress reached OpenAI, that OpenAI then put a huge amount of compute on a similar and uncommon route, and that conversations about credit followed. TechCrunch and WIRED have the reported timeline.

OpenAI’s own write-up, On the Navier–Stokes Millennium Prize Problem, says its researchers and agents did not see the pair’s work before it was public. In the same breath, the company says it cannot rule out that de-identified data from product usage helped its models.

I am not here to pick a courtroom winner. I am here because that second sentence exists. The operator of the product is telling you, in public, that work you type into the product might still make the model better—the same model that can then compete with you, or be sold back to you.

If that does not make you ask who owns the drafts you paste into ChatGPT, Codex, Claude, or Claude Code, it should.

Read the fine print

Most people treat those boxes like a notebook. They are not. They are a service. The history lives on the company’s servers. The terms decide what the company may do with it. For a plain-language map of the copies—history, training, human review, legal records—see What happens to what you type?.

As of this writing, the big consumer apps usually train on you unless you opt out. ChatGPT’s “improve the model” style settings, Claude’s help-improve toggles, Gemini’s activity controls, Codex’s training defaults: the burden is on you to find the switch. Flip it and you are still trusting the operator to honor it. Paid workplace SKUs often default the other way, which is the tell. Companies with lawyers negotiated “do not train on us.” Individuals got a checkbox buried in settings.

Even then, you are trusting them. Models need data to get better. The easiest, cheapest, most current source of that data is you—your proofs, your code, your emails, your half-finished ideas. They improve the model on that material and sell the improved model back. That is not a conspiracy. It is the incentive.

A no-training toggle does not change the incentive. It asks you to believe the incentive lost.

Why we built Ethyx

I did not want the only option to be “park your entire working life inside one provider’s consumer app and hope the fine print is kind.”

Self-hosted clients already exist. LibreChat is the one people usually mean, and it is a good one. If you can run your own stack, you should consider it. Keeping history on hardware you control is, for a lot of cases, a stronger privacy posture than any managed app—ours included. (If you still point that stack at OpenAI or Anthropic, the live prompt still leaves the house. The win is that the archive is yours.) It also asks you for API keys, updates, backups, and the patience to fix it when it breaks. Most people should not have to become a sysadmin to have a right to privacy. That is the other reason we started Ethyx.

So we built a client where your history lives with Ethyx. Conversations, titles, project notes, uploads: stored here, encrypted at rest. You are not accumulating a second corpus of your life inside ChatGPT or Claude that those companies keep under their consumer terms.

When a turn runs, the live prompt still has to reach a model. That part cannot be wished away—we explained why in Why Ethyx isn't end-to-end encrypted. What we can change is how it gets there:

  • We call commercial APIs, not the consumer ChatGPT / Claude / Gemini apps. API terms are a different bargain than “help improve the product for everyone.”
  • Where the API supports it, we send a no-retention flag (store: false) so the provider is asked not to keep an application copy of the turn. Anthropic’s commercial API, as published, does not train on API conversations by default.
  • We pseudonymize your identity before the request leaves, so the provider is not also collecting your name and email with the prompt.

That is a smaller, clearer bargain than handing a consumer app your whole archive. It is not a disappearing act.

Be fair: you still trust the provider to abide by those API terms. A flag is a request, not a guarantee. We do not have zero-data-retention contracts. We do not control their infrastructure. Anyone who tells you otherwise is selling comfort. We are trying to improve the process as far as the actual plumbing allows—and to say so without a slogan.

Ethyx would not have stopped two labs from racing the same math problem after a rumor. That is not the claim. The claim is that you should not have to put every unpublished draft, every coding session, and every private thread into a product whose operator can say, out loud, that your usage might still have helped the model.

Why you ought to care

Care who owns the work you give an AI. Care where it is stored. Care what can be done with it after you hit send.

This week’s drama is a loud version of a quiet default. Most people will never prove a Millennium problem. Plenty of people will paste a manuscript, a codebase, a medical question, or a business plan into a box that feels private and is not.

If that trade-off bothers you, that is why Ethyx exists. We are in closed beta. Join the waitlist below, or contact us. I read those.

Get early access

Ethyx is in closed testing. Leave your email and we will let you know when access opens up, plus an occasional note on what shipped.

We email you about access and product updates only. Unsubscribe any time. Privacy policy