Alex Wang’s Post and the New Data Scarcity
данные для ИИсинтетические данныеScale AI
What can be stated with confidence about the post
I would not pretend to know what the post said: the exact text of Alex Wang’s X publication is unavailable in the materials provided. The indexed entry is dated October 2, 2026, so attributing specific wording or conclusions to it would be fabrication.
The available context is still quite concrete. In related interviews and Scale AI materials, Wang has consistently described a shift away from large-scale data collected from the open internet toward proprietary enterprise datasets, expert-created data for advanced tasks, and synthetic examples reviewed by humans.
Technically, this means the bottleneck has changed. It is not enough to generate another large set of plausible answers: teams need selection rules, domain specialists, error controls, and an evaluation loop that does not reward a model for confidently packaged nonsense.
Hybrid human-AI pipelines stand out as a central idea. Models help generate and scale data, while people guide the process and validate quality. This does not eliminate annotation; it moves it to a more demanding level, from simple labels to judging reasoning, specialized answers, and system behavior.
Why model selection no longer solves the problem
The practical implication is stronger than another debate about models: durable advantage is shifting toward data and feedback. The same foundation model can deliver fundamentally different results depending on corpus quality, evaluation design, and how errors are fed back into the pipeline.
For developers, that changes the order of priorities. I would first examine data provenance, coverage of rare cases, consistency among expert judgments, and safeguards against accumulating synthetic errors. Only then does it make sense to discuss fine-tuning or replacing the model.
For Scale AI, this argument naturally expands the company’s role from annotation toward data and evaluation infrastructure. Yet the core engineering tension remains unchanged: synthetic data is easy to scale, while trust in that data is still difficult to scale.