The greenest compute is the compute you already own
Every watt a program wastes is a watt someone generated. Artemis makes existing code measurably more efficient — so the same work takes less energy, the same hardware serves for longer, and the research that matters runs sooner.
- Less energy
- per result, measured on real hardware
- Longer life
- for the hardware the world already owns
- Faster science
- research code that runs in hours, not weeks
Why optimization is a public good.
Software efficiency compounds: one merged improvement runs on every machine, every day, for every user downstream.
Energy that never gets burned
Optimized code does the same work with fewer joules. When an inference stack serves 40% more requests on the same machines, that is energy nobody has to generate — at data-center scale, every percentage point matters.
✓ 4.9× kernel speed-up merged into OpenAI's Whisper
Hardware that lives longer
The most wasteful upgrade is the one you didn't need. Artemis recovers headroom from the machines you already own — CPUs serving models that supposedly needed GPUs, clusters at 95% capacity given room to breathe — deferring procurement and e-waste alike.
✓ 2.07× LLM throughput on existing Intel CPUs — an approved, official build
Science that runs sooner
Research code is written to answer questions, not to be fast — and slow code quietly rations science. From risk models tuned for accuracy to simulation loops cut from weeks to days, optimization gives researchers their iteration speed back.
✓ Production ML models optimized for accuracy, not just speed
Not just faster. Better.
Performance is one metric among many — Artemis optimizes whatever a project can measure, and for research that often means accuracy.
ML models that predict better
Artemis tunes models against the metric that matters — accuracy. A vision model that misses less, a cancer-detection model that catches more, a risk model that errs more rarely: if your research measures it, Artemis can search for it. We have already optimized production ML models for accuracy at a major UK bank.
Agents that think for less
Our Evolving Excellence paper shows Artemis jointly tuning the prompts, tools, and parameters of LLM-based agents — same answers, around 37% fewer tokens, with accuracy gains across programming and mathematical-reasoning workloads. No architecture change required.
Read the paperYour research metric is the goal
Speed, energy, memory, accuracy, recall on the class you care about — Artemis is metric-agnostic by design. Many researchers have already built their thesis work on it, using the same engine that optimizes production systems to push their own results further.
Work on something that matters? Send it to us.
If you maintain an open-source project or run a research codebase — climate, health, science, infrastructure — submit the repo or describe the problem, whether the metric is speed, energy, or accuracy. We'll apply Artemis to it and, where the results hold up, contribute them back upstream as pull requests you can inspect line by line.
- ggml-org/llama.cpp+13% tokens/s
- ggml-org/llama.cpp≈1% tokens/s
- ggml-org/llama.cpp5.2× and 47× on the copy kernels
- QuentinFuxa/WhisperLiveKit≈½ the encoder calls
Send us the project.
Climate, health, science, infrastructure — or the code behind a paper. Tell us what it does and what matters: speed, energy, or accuracy. No repo yet is fine; describe the problem.
- 24hr reply
- 10landed upstream
- 0cost to you
People on the problems, not just software.
TurinTech runs a standing research-intern program, and many researchers have already built their thesis work on Artemis, supervised by our engineers — our NVIDIA result was proposed by an intern in their first week. For projects that matter, we can put TurinTech engineers and students to work on your problem directly and contribute what they find.