
Define better, in your metrics
Set the goal and the baseline (latency, throughput, cost, accuracy). Every version that follows is measured against it.
Describe a goal and the agent runs the full cycle autonomously.
Make the two hot attention kernels faster on CPU, without changing what they return.
Both run as fragmented eager ops today — too much memory traffic on decode, and the prefill loop is still in fp32.






