All publications

Artemis AI: Multi-LLM Framework for Code Optimisation

2025·IEEE CAI 2025
LLM₁LLM₂LLM₃populationshared candidatesbenchmarkreal hardwareminimal diffships as a PRcollaborative generation, one selection pressure

Figure 1: Collaborative multi-LLM optimization. Generator models contribute candidates to one population (a); hardware benchmarking selects survivors (b); the winning variant ships as a minimal diff (c).

The platform paper. Artemis coordinates multiple LLMs as variant generators over a shared candidate population, with benchmarking on real hardware as the selection pressure. Collaboration beats any single model: different models contribute different optimization ideas, and the population keeps whichever survives measurement.

Across domains the framework delivers substantial execution-time reductions with deliberately minimal diffs — small changes are easier to review, merge, and maintain, which is what makes the results usable in production rather than only in papers.

Key results

  • Substantial execution-time reductions across application domains
  • Minimal-diff changes designed for real review and merge
  • Published at IEEE Conference on Artificial Intelligence 2025