-
Notifications
You must be signed in to change notification settings - Fork 2.1k
Pull requests: JustVugg/colibri
Author
Label
Projects
Milestones
Reviews
Assignee
Sort
Pull requests list
olmoe: decode goes disk-bound whenever the expert LRU misses
#698
opened Jul 29, 2026 by
ariannamethod
Loading…
fix(cuda): evict cold experts and retry instead of losing a resident tensor (#687)
#696
opened Jul 29, 2026 by
terrizoaguimor
Loading…
5 tasks done
feat(cuda): say it once, loudly, when resident tensors fall back to CPU (#687)
#694
opened Jul 29, 2026 by
terrizoaguimor
Loading…
3 tasks done
Kimi K3: integrate coli chat, API, and Web
#684
opened Jul 29, 2026 by
ZacharyZcR
Contributor
Loading…
Kimi K3 engine: native-MXFP4 expert streaming, KDA + gated NoPE MLA + AttnRes + LatentMoE (#658)
enhancement
New feature or request
model-support
Supporto a nuovi modelli
#676
opened Jul 28, 2026 by
steve-m
Contributor
Loading…
feat: add measured machine autotuning
enhancement
New feature or request
needs-rebase
Confligge, serve rebase dell'autore
performance
Velocità / tok-s / ottimizzazioni
#673
opened Jul 28, 2026 by
ZacharyZcR
Contributor
Loading…
feat(win): fix silent CPU fallback, launcher suite, DirectStorage expert loads
enhancement
New feature or request
needs-rebase
Confligge, serve rebase dell'autore
#670
opened Jul 28, 2026 by
khalilswdp
Loading…
8 tasks done
feat(core): model-architecture seam, chat templates, text antiprompt
enhancement
New feature or request
needs-rebase
Confligge, serve rebase dell'autore
#667
opened Jul 28, 2026 by
khalilswdp
Loading…
5 tasks done
cuda: add experimental lossless compressed expert tier
cuda
Backend CUDA/NVIDIA
enhancement
New feature or request
#651
opened Jul 27, 2026 by
ZacharyZcR
Contributor
Loading…
tools: deprecate convert_olmoe.py
enhancement
New feature or request
needs-rebase
Confligge, serve rebase dell'autore
#650
opened Jul 27, 2026 by
CodePrometheus
Loading…
3 of 5 tasks
QLoRA training path: fine-tune GLM-5.2 (744B) in 64 GB RAM
enhancement
New feature or request
feature
Nuova funzionalità
#626
opened Jul 26, 2026 by
pavolbauer
Loading…
4 of 5 tasks
Add opt-in per-expert causal-ablation harness (ABLATE_SCORE) for moe()
enhancement
New feature or request
#621
opened Jul 25, 2026 by
jeswr
Loading…
docs(api): OpenAI server supports tool-calling
docs
Documentazione
needs-rebase
Confligge, serve rebase dell'autore
#620
opened Jul 25, 2026 by
jeswr
Loading…
Add Qwen3.6-35B-A3B engine: Vulkan MoE backend, resident expert pinning, serve GPU support + fixes
model-support
Supporto a nuovi modelli
needs-rebase
Confligge, serve rebase dell'autore
vulkan
Backend Vulkan/AMD
#602
opened Jul 25, 2026 by
minne100
Loading…
4 of 5 tasks
WIP: MiniMax-M3 support — GQA + MSA block-sparse attention, o200k tokenizer, converter (follow-up to #418)
model-support
Supporto a nuovi modelli
#601
opened Jul 24, 2026 by
steve-m
Contributor
Loading…
5 tasks
Metal fmt=4 grouped-int4 decode: attention + routed experts (#585)
metal
Backend Metal/Apple
#587
opened Jul 24, 2026 by
RDouglasSharp
Contributor
Loading…
Fix stateful KV tail at the NGEN limit
bug
Difetto verificato nel codice
#567
opened Jul 23, 2026 by
winklemad
Contributor
Loading…
3 of 5 tasks
add persistence controls for lower SSD writes
enhancement
New feature or request
#555
opened Jul 23, 2026 by
Skater1808
Loading…
5 tasks
CPU: KV cache quantization — KV8 (fp8 e4m3) + KV_TQ (rotated-int4 / PolarQuant)
enhancement
New feature or request
performance
Velocità / tok-s / ottimizzazioni
#553
opened Jul 23, 2026 by
NeuralNotwerk
Contributor
Loading…
feat: add WebGPU expert workers
enhancement
New feature or request
#552
opened Jul 23, 2026 by
gauravsaini
Loading…
feat: add distributed expert workers
enhancement
New feature or request
#551
opened Jul 23, 2026 by
gauravsaini
Loading…
feat: add dense MLP activation sharding
enhancement
New feature or request
performance
Velocità / tok-s / ottimizzazioni
#550
opened Jul 23, 2026 by
gauravsaini
Loading…
feat: add Qwen3-30B-A3B engine
model-support
Supporto a nuovi modelli
#544
opened Jul 23, 2026 by
opxyc
Contributor
Loading…
4 of 5 tasks
feat: interactive chat mode for olmoe (multi-turn conversations)
enhancement
New feature or request
#542
opened Jul 23, 2026 by
opxyc
Contributor
Loading…
4 of 5 tasks
Previous Next
ProTip!
Filter pull requests by the default branch with base:main.