Local Inference: 90W vs Expensive GPUs
локальный инференсоблачный инференсGPU
Local deployment saves money in unexpected places
I would not reduce the decision to token pricing alone. In the example discussed, a local system consumed around 90W and made it possible to experiment without worrying about an unexpected cloud bill. In the original discussion, participant 144406 cited Spark as an example of such a setup.
Still, 90W should not be treated as a universal benchmark for local AI. It reflects a specific operating mode, model, and hardware configuration—not a guideline for serving large models. As memory and performance requirements grow, a local setup can quickly shift from an efficient workstation into an expensive rack that also needs cooling.
This is where the open-weight ecosystem becomes paradoxical. Near-daily releases of Chinese models broaden the available choices, while simultaneously driving demand for GPUs with large memory capacity. At the time of the discussion, RTX PRO 6000 pricing had risen from $11–13K to $16–17K.
Cloud services remove the upfront investment, but every active session has a cost. Mid-2026 pricing snapshots placed H200 instances at roughly $4.50–5.50 per hour on standard rates and $1.80–2.80 per hour through some spot offers. For infrequent, demanding runs, that arithmetic can be more compelling than buying an accelerator.
The contrast is clear in Anthropic's August 28, 2026 article, “Automated researchers can reliably mitigate alignment failures.” Its automated researchers used H200 hardware—a class of workload for which a compact 90W local setup is not a direct substitute.
Utilization, not ideology, now determines the choice
Local inference wins when workloads are steady, data is sensitive, and full control over the environment matters. The cloud remains stronger for irregular demand, sudden peaks, and experiments that temporarily require H200-class accelerators.
The first calculation should include more than electricity and an hourly rate. A local cost model includes GPU purchase price, idle time, cooling, maintenance, and the risk of rapid obsolescence. A cloud model includes storage, data transfer, and hours lost to poorly managed resource lifecycles.
Rising RTX PRO 6000 prices show that open-weight models do not merely democratize inference. They shift scarcity from APIs to hardware. The central unresolved question is no longer where the model runs, but how fully the accelerator you buy will actually be utilized.