pull down to refresh

Interesting, thanks. Been wanting to try llama_cpp but didn't have the hardware at the time to run any kind of productive models. Not sure I do now either, but they seem to compress better and better.

What factors made you choose the approach you did?

I won't pretend I'm a dev, so maybe I could use some sort of framework to start with, unless the docker container you linked does just that?