Artemis AI: Multi-LLM Framework for Code Optimisation
Figure 1: Collaborative multi-LLM optimization. Generator models contribute candidates to one population (a); hardware benchmarking selects survivors (b); the winning variant ships as a minimal diff (c).∎
The platform paper. Artemis coordinates multiple LLMs as variant generators over a shared candidate population, with benchmarking on real hardware as the selection pressure. Collaboration beats any single model: different models contribute different optimization ideas, and the population keeps whichever survives measurement.
Across domains the framework delivers substantial execution-time reductions with deliberately minimal diffs — small changes are easier to review, merge, and maintain, which is what makes the results usable in production rather than only in papers.
Key results
- Substantial execution-time reductions across application domains
- Minimal-diff changes designed for real review and merge
- Published at IEEE Conference on Artificial Intelligence 2025