is the 4k context length of llama2 for real?

actually-a-cat@sh.itjust.works · edit-2 1 year ago

You are supposed to manually set scale to 1.0 and base to 10000 when using llama 2 with 4096 context. The automatic scaling assumes the model was trained for 2048. Though as I say in the OP, that still doesn’t work, at least with this particular fine tune.

actually-a-cat@sh.itjust.works · 1 year ago

is the 4k context length of llama2 for real?

actually-a-cat@sh.itjust.works · 1 year ago

Those are OpenCL platform and device identifiers, you can use clinfo to find out which numbers are what on your system.

Also note that if you’re building kobold.cpp yourself, you need to build with LLAMA_CLBLAST=1 for OpenCL support to exist in the first place. Or LLAMA_CUBLAS for CUDA.

actually-a-cat@sh.itjust.works · 1 year ago

What’s the problem you’re having with kobold? It doesn’t really require any setup. Download the exe, click on it, select model in the window, click launch. The webui should open in your default browser.

actually-a-cat@sh.itjust.works · edit-2 1 year ago

Small update, take what I said about the breakage at 6000 tokens with a pinch of salt, testing is complicated by something somewhere breaking in a way that persists through generations and even kobold.cpp restarts… Must be some driver issue with CUDA because it takes a PC reboot to resolve, then the exact same generation goes from gibberish to correct.

actually-a-cat@sh.itjust.works · edit-2 1 year ago

kobold.cpp now supports NTK scaling and it works

actually-a-cat@sh.itjust.works · 1 year ago

Why is the front page suddenly so stale?

actually-a-cat@sh.itjust.works · 1 year ago

Reddit has over 2,000 employees most of whom are doing bullshit nobody using the site actually needs or wants, it’s possible to run a lot leaner than that. Like Reddit itself used to, before they started burning hundreds of millions trying to compete with every other social media site at once instead of being Reddit

actually-a-cat

is the 4k context length of llama2 for real?

is the 4k context length of llama2 for real?

kobold.cpp now supports NTK scaling and it works

kobold.cpp now supports NTK scaling and it works

Why is the front page suddenly so stale?

Why is the front page suddenly so stale?