Inference
When Does a Second GPU Actually Help LLM Inference?
095
There is an appealingly simple upgrade path for a busy local AI server: install another GPU. Twice the accelerators should mean something close to twice