Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

You’re a mad man - thank you!

Do I understand correctly that Ollama doesnt do that, and that’s why responses hang forever on a M3 running the same model through Ollama?




Doesn't Ollama use llama.cpp so their point stands even if they used it directly?


No, because Ollama is buggy. The first step to answering GPC’s question is to try using an up to date llama.cpp.


It wouldn't be the first time ollama's llama.cpp fork reintroduced bugs and was missing important optimizations.


lol jealous hater


Huh?


Thank you!

afaik ollama relies on llama.cpp and mmap. mmap loads pages on demand and doesn't use the same explicit cache or parallel reads like my engine. Most likely ollama/llama.cpp will be way slower in this case




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: