Research
Building Better AI by Improving How It Learns
Building Better AI by Improving How It Learns
Project: WILDCHAT-50M: Understanding the Role of Synthetic Data in AI Training
Lead researcher: Chinmay Hegde, Professor of Computer Science and Engineering, New York University
Institution: New York University
Large language models learn general knowledge during their initial training, but they become more useful through a second stage known as post-training. During this process, AI models learn to follow instructions, respond more accurately, and interact more naturally with people.
Professor Chinmay Hegde and his team at New York University are studying one of the most important ingredients in post-training: the quality of the data used to teach AI models after their initial training.
Their project created WILDCHAT-50M, the largest publicly available dataset of AI-generated conversations. By comparing responses from more than 50 different open-source language models, the researchers are identifying which kinds of synthetic training data produce the strongest and most reliable AI systems. The project also introduced a new training approach, called RE-WILD, that outperformed several leading alternatives while using significantly less training data.
The challenge
Today’s most advanced AI systems depend heavily on synthetic data—responses, examples and evaluations generated by other AI models rather than written directly by people.
Much of this data is created inside private technology companies and is not publicly available. As a result, university researchers often lack access to the large, high-quality datasets needed to study how AI models learn most effectively.
This makes it difficult to answer important questions. Does a larger AI model always generate better training data? How much data is actually needed? Does combining responses from many models improve performance? Which characteristics of a model are passed on during training?
Without standardized public datasets, researchers have had few opportunities to compare these approaches at scale, slowing progress in open AI research.
A new approach
To address this challenge, the NYU team created WILDCHAT-50M by expanding an existing collection of real-world user prompts with responses generated by more than 50 open-weight language models ranging from 500 million to 104 billion parameters.
Each model participated in more than one million multi-turn conversations, creating approximately 125 million conversational exchanges—more than 50 times larger than the next-largest comparable public dataset identified by the researchers.
The researchers then used this resource to study how different synthetic datasets affect the performance of language models during post-training.
To demonstrate the value of the dataset, they developed RE-WILD, a carefully designed training mixture that achieved stronger performance than several leading public alternatives while using only about 40 percent as many training examples.
How Empire AI makes this possible
Creating and analyzing WILDCHAT-50M required an enormous amount of computing power.
The project generated responses from more than 50 language models over roughly two months using approximately 10,000 NVIDIA H100 GPU hours. Researchers also trained and evaluated numerous AI models to measure how different synthetic datasets affected performance across multiple benchmarks.
The paper acknowledges computing support from the Empire AI Consortium, along with resources from NYU and the National Artificial Intelligence Research Resource Pilot. Access to high-performance computing allowed the team to conduct experiments that would have been impractical for most academic laboratories.
Empire AI helps researchers perform controlled comparisons at a scale needed to understand how AI systems learn, helping narrow the gap between academic research and the capabilities of large commercial AI developers.
Key findings
The research showed that the source of synthetic training data has a surprisingly large influence on how well an AI model performs after training.
Among the team’s findings:
- Bigger AI models do not always generate better training data.
- Increasing the amount of high-quality synthetic data generally improves performance.
- Carefully selecting one strong source of training data often works better than combining responses from many different models.
- AI models inherit writing style, organization and clarity from the models that generate their training data more consistently than they inherit factual knowledge or mathematical ability.
- Different language models often produce much more similar responses than researchers expected.
These findings suggest that improving AI is not simply about collecting larger datasets. Choosing better training data may be just as important as increasing its quantity.
Potential impact
By making WILDCHAT-50M publicly available, the researchers are giving universities, startups and nonprofit organizations a valuable resource for developing stronger AI systems without relying exclusively on proprietary data.
The project could help researchers:
- Build more capable open-source language models.
- Reduce the cost and time required for AI post-training.
- Better understand which synthetic data leads to the strongest AI performance.
- Improve the reproducibility of AI research.
- Advance safer and more reliable AI systems through better training methods.
Rather than treating synthetic data as a black box, Professor Hegde’s team is helping establish a scientific foundation for understanding how AI learns after its initial training. Their work provides researchers with both the data and the evidence needed to develop the next generation of open, transparent and trustworthy AI systems.
For New York: Because the dataset is publicly available, researchers across New York can build on this work to develop new AI applications in healthcare, finance, education, cybersecurity and scientific research—industries that play a major role in the state’s economy.”
Note: The paper on this research was accepted for presentation at the International Conference in Machine Learning in 2025.