Hugging Face integra kernels do llama.cpp ao transformers: GGUF carrega com from_pretrained, roda em Mac e mantém toda a API do PyTorch. Veja o trade-off.
Some results have been hidden because they may be inaccessible to you
Show inaccessible results