Performance
A local LLM does not have one meaningful “tokens per second” number. Before it can write the first token of an answer, it has to process the tokens already in the prompt.
Local AI, Local LLMs, GPU hardware, inference, RAG and AI server infrastructure explained clearly.