Inference
When Does a Second GPU Actually Help LLM Inference?
094
There is an appealingly simple upgrade path for a busy local AI server: install another GPU. Twice the accelerators should mean something close to twice