Data4Scientists aims to enable scientists around the world to regain control over their research data used by Artificial Intelligence (AI) and better defend their rights.
By reverse-engineering AI algorithmic systems using the DIGIPOWER tool and leveraging the right of access to data, the Data4Scientists project aims to help scientists (mathematicians, scientists, engineers, etc.) to gather evidence of AI systems’ use of their data and thereby assert their intellectual property rights.
This is another iteration of our Data4All project, following Data4Workers ⤤ and Data4Mods ⤤
BACKGROUND AND ORIGINS OF THE PROJECT
On September 8, 2026, OpenAI announced that it had found a solution to the existence and regularity problem for the Navier-Stokes equations, one of the Millennium Prize Problems set by the Clay Mathematics Institute.
According to OpenAI, the large-scale deployment of AI agents made it possible to obtain a solution to the Navier-Stokes problem. This claimed result is potentially of historic mathematical significance, but—as is usually the case with a result of this magnitude—it remains subject to thorough scrutiny by the mathematical community.
This mathematical announcement was accompanied by a second controversy. Mathematicians Tristan Buckmaster and Levent Alpöge (an employee at Anthropic) were working independently on closely related problems involving fluid equations, using AI systems as part of their research. Their work included a result concerning Euler’s equations and related advances toward the Navier-Stokes equations. OpenAI stated that reports suggesting another group might be on the verge of solving a “millennium problem” had prompted it to focus exceptionally large computational resources on the Navier-Stokes equations. Buckmaster subsequently raised questions about the relationship between these two research efforts, the issue of attribution, and the possibility that interactions with OpenAI’s systems might have indirectly influenced OpenAI’s models. OpenAI disputes key elements of his account and asserts that it did not use the researchers’ instructions or demonstrations to guide the agents that produced this result.
Since then, other mathematicians ⤤ who have found themselves in similar situations over the past year have begun to speak out.
One particularly important question remains unanswered. OpenAI has stated that, while it considers this unlikely, it cannot categorically rule out the possibility that de-identified information derived from researchers’ use of its products may have contributed to improving the models in question. It has since reinforced this statement regarding user data collected over the past few weeks.
These questions are very different from those concerning the mathematical correctness of the proof. They concern the social and institutional conditions under which AI-assisted mathematics is emerging.
THREE SEPARATE QUESTIONS
1. What happened in this specific case? This is a factual question, certain aspects of which remain disputed.
2. What should be considered a scientific contribution when research is mediated by AI? This is essentially a question for the scientific communities themselves to address.
3. What procedural and institutional conditions should govern AI systems that are playing an increasingly significant role in scientific research? This is a question on which scientific communities, learned societies, technology providers, universities, governments, and human rights organizations all have a say.
how to (re)act
This controversy offers an exceptionally concrete opportunity for mathematicians, researchers, and human rights experts to examine together a series of questions that are bound to become increasingly important: What does it mean to “participate in science” when artificial intelligence systems serve as research tools, repositories of scientific interactions, and increasingly capable actors that are integral to scientific discovery?
get involved
Are you a scientist interested in participating in the project? Contact us so we can help you recover, analyze, and protect your data, and thus build a community capable of collectively regaining control over the systems they help power.