Posts
All the articles I've posted.
Getting Fill-In-the-Middle Autocomplete Working in VS Code Continue with llama.cpp
Assign a dedicated autocomplete model to llama-server in Continue's YAML configuration, tune suggestion latency, and keep chat endpoints separate when troubleshooting FIM completion.
Notes on Tabby: Llama.cpp, Model Caching, and Access Tokens
Notes on Tabby's local inference backend, configurable model cache directory, model registry, and browser-based account setup for obtaining an access token.
Setting Reasoning Strength in OpenWebUI with `chat_template_kwargs`
Pass reasoning_effort through an OpenWebUI custom model parameter to a llama.cpp chat template, with examples of low, medium, and high settings and their tradeoffs.
A Simple Tool for Downloading Files from Hugging Face
A Python downloader accepts Hugging Face URLs or repository paths, preserves revisions, previews downloads, and falls back from individual files to directory snapshots.