Median time to first token and time per output token, from 60 to 90 benchmark prompts per cell (GSM8K, IFEval, BFCL). Prompts run 19 to 772 tokens, median 92.
Each line joins one chip's two runtimes. Lower and left is faster. Both axes are logarithmic.
The same numbers as a grid, both columns on logarithmic scales. llama.cpp ExecuTorch. The model picked above is banded.