Intel has published a partner brief on TurinTech Artemis and added Artemis as an official optimisation platform.
The brief follows a benchmark Intel ran and formally approved. Artemis optimised vLLM CPU inference for Qwen3.6-35B-A3B and delivered 2.07× output token throughput — from 38.17 to 79.32 tokens per second at BF16, concurrency 8, on an Intel Xeon 6767P (Granite Rapids, 128 CPUs).
What matters about that number is where it comes from. It is not a smaller model or a lower precision standing in for the real thing: same model, same weights, same output. The gain is in the code around the model, which is the only kind of speedup a risk committee never has to argue about.
The relationship has since widened past the benchmark. An Intel engineer picked up one of our GPU kernel optimisations and raised it upstream in OpenVINO themselves, and through Intel we are now working with Phison on their AI acceleration library.



