An optimized quantization and inference library for running LLMs locally on modern consumer-class GPUs