Gorgon Halo Benchmarks
Nvidia is expected to launch RTX Spark devices later this week, following a tease at the end of last week and the chip’s announced fall release. Ahead of Microsoft’s event on Wednesday, October 7, AMD has shared some benchmarks for its new Gorgon Halo chips, the range that will directly compete with RTX Spark devices. The extra insight comes a matter of days after the first Gorgon Halo devices launched, some of which cost upwards of $7,099.
Short of a few questionable Geekbench leaks, performance results for the RTX Spark have not yet been widely seen, so AMD is not using it as a direct comparison point. Rather, the company is comparing the Ryzen AI Max+ Pro 495 to Intel’s Core Ultra X9 388H. These chips are not in the exact same class of device, though AMD argues it is comparing its top-of-stack part to Intel’s top-of-stack part.
AMD used ComfyUI to measure generative AI performance. The company averaged multiple runs of various models, comparing total throughput to Intel’s competition. AMD used its top-spec 192GB configuration of the Ryzen AI Max+ Pro 495 and compared it to the Core Ultra X9 388H in a system with 64GB of memory.
The performance advantage ranges from 1.1x up to 32.2x, though the end point is an outlier. It is possible this delta comes down to an optimization issue, or that the workload is simply too large to run effectively on the competing machine.
AI Performance and Memory Capacity
General application and gaming performance figures remain sparse, but major shifts compared to last-generation Strix Halo chips are not expected, as Gorgon Halo functions largely as a refresh of that range.
Gorgon Halo targets the RTX Spark segment and, to a lesser extent, Apple’s larger M-series SoCs. This category of agentic PCs is expanding, though its market footprint remains smaller than initial hype suggests. AMD noted in press briefings that it has shipped over half a million agentic PCs to date, referring primarily to Strix and Gorgon Halo hardware.
The matchup between Gorgon Halo and the RTX Spark focuses heavily on memory capacity. RTX Spark devices top out at 128GB of unified memory, similar to the GB10 in the DGX Spark. AMD, by contrast, supports up to 192GB with Gorgon Halo, enabling larger models to run locally, albeit potentially at a lower performance level.
When running the GLM 5.3 Flash model with 320 billion parameters, AMD recorded a peak throughput of 20 tokens per second using Unsloth's UD-IQ4_XS mixed-quantization format. By comparison, previous tests on the DGX Spark running GPT-OSS 120B with 4-bit quantization clocked 64 tokens per second, while the last-gen Ryzen AI Max+ 395 reached 56 tokens per second.
Although GLM 5.3 Flash features 320 billion total parameters, only 18 billion are activated for each token. Similarly, GPT-OSS 120B is a mixture-of-experts model with 120 billion total parameters, with roughly 5 billion active per token.
AMD also outlined performance for Qwen 3.8 Flash Next, a multimodal MoE model featuring 125 billion main parameters, 51 billion embedding parameters, and about 4 billion parameters for multi-token prediction. According to AMD, the Ryzen AI Max+ Pro 495 reaches up to 42 tokens per second with this model using dynamic 4-bit quantization and MTP.
While these figures are competitive, performance tends to drop as context lengths increase during local execution. If initial throughput for GLM 5.3 Flash starts at 20 tokens per second, extended context lengths could introduce bottlenecks for larger models.
Hardware Specifications
Gorgon Halo represents a spec refresh of Strix Halo, featuring an increase to 192GB of unified memory from 128GB and a 100 MHz boost clock increase for the 495 model. The rest of the product stack retains identical core counts and microarchitectures.
Testing with proxy hardware indicates that while additional unified memory permits the execution of larger models, higher capacity does not automatically translate to faster inference speeds.
Initial Gorgon Halo devices, such as the Minisforum MS-S1 Max-P495, are currently available. Top-line configurations are priced around $7,000, with broader market pricing expected to span a wide range above that threshold.




