First run of Mistral3 8B Q4 model on Steam Machine — ~700 tok/s prompt processing and ~43 tok/s token generation using llama.cpp.
Going to put llama-swap on there to try out a couple of different models…which is written by my old friend Benson @mostlygeek.bsky.social!