3 min read

AMD MI350P: 144GB of Memory and 4TB/s Bandwidth

AMD Instinct MI350PHBM3Eинференс LLM

AMD Instinct MI350P combines 144GB of HBM3E memory with up to 4TB/s of memory bandwidth. That makes it a notable PCIe option for large-model inference, where fitting weights and context on one accelerator matters. AMD has not published an official price, so estimates near $20,000 remain speculative.

What AMD MI350P actually offers

The headline figure comes first: AMD Instinct MI350P features 144GB of HBM3E memory with peak bandwidth of 4TB/s. AMD lists these specifications, along with a 4096-bit memory interface, on the product page. For a single PCIe accelerator, that is a serious proposition for large-model inference.

Memory capacity matters here more than an abstract comparison of compute units. With 144GB, there is a better chance of keeping a model, an expanded context window, or a larger request batch on one device. That does not guarantee fast generation, but it reduces the need to split a model across multiple accelerators.

The 4TB/s bandwidth figure is especially relevant for operations constrained by reading weights from memory. Still, it is a peak specification, not a promised speed for a particular LLM. Real performance will depend on weight format, batch size, context length, inference kernels, and the maturity of the software stack.

MI350P belongs to the ROCm ecosystem and the CDNA 4 architectural line with AMD Infinity. AMD positions it as a server accelerator for standard air-cooled systems, not as a workstation graphics card. Comparing it directly with an RTX 6000 based on memory alone may look compelling, but it mixes different device classes.

Pricing is less clear. As of August 20, 2026, AMD had not listed an official retail price on the publicly available product page. The estimate of roughly $20,000 appears to be a market assumption, so it is too early to base infrastructure cost calculations on it.

Why 144GB really changes the equation

For inference, this is not a cosmetic update: extra memory can replace a complicated multi-card setup with a single accelerator. Less device-to-device communication means simpler deployment and potentially more predictable latency, although the final result must still be validated on the specific model.

The biggest gain will come in scenarios where a model nearly fits into a smaller memory pool, or where context and batching quickly consume the remaining headroom. Capacity alone does not remove ROCm limitations, framework compatibility concerns, or the need for high-quality optimized kernels. I would first verify support for the required model and the stability of its operating profile, then look at the attractive terabytes-per-second figure.

MI350P looks like a credible alternative to NVIDIA accelerators for large-scale inference, but not yet an obvious replacement. The deciding factors will be more than its 144GB capacity: availability, actual price, and how easily these specifications translate into tokens per second.

Previously, we examined Cocoon and how confidential computing changes the economics and privacy of AI inference. AMD’s new HBM3E capacity and bandwidth figures show why infrastructure for these workloads is becoming a competitive battleground.