Speeding up local LLMs has moved beyond the competition of just making models smaller, into a phase that combines lookahead ...