-
2Wild Agency
- Atlanta, GA
-
15:12
(UTC -12:00) - https://youtube.com/@Tonyd2wild
- @2WildTech
- @Tonyd2wild
- @Tonyd2wild
Popular repositories Loading
-
DeepSeek-v4-Flash-DSpark-1M-NVFP4-KV-2x-DGX-Spark
DeepSeek-v4-Flash-DSpark-1M-NVFP4-KV-2x-DGX-Spark PublicDeepSeek V4 Flash DSpark 1M NVFP4 KV recipe for 2x DGX Spark
-
GLM-5.2-QuantTrio-200K-4x-DGX-Spark--36tok-s
GLM-5.2-QuantTrio-200K-4x-DGX-Spark--36tok-s PublicRecipe: GLM-5.2 (unpruned QuantTrio Int4-Int8Mix) at 200K ctx with MTP spec decode on a 4x NVIDIA DGX Spark (GB10) cluster
-
Deepseek-v4-Flash-TP2-DGX-Spark-500k-CTX
Deepseek-v4-Flash-TP2-DGX-Spark-500k-CTX PublicWorking recipe to serve DeepSeek-V4-Flash across two NVIDIA DGX Spark (GB10) nodes with vLLM (TP=2, FP8 KV, MTP) over a RoCE/RDMA link — Docker image, launch scripts, RDMA/NCCL setup, and the gotchas.
-
MiniMax-M3-2x-DGX-Spark-36-tok-s
MiniMax-M3-2x-DGX-Spark-36-tok-s PublicMiniMax-M3 (428B, no pruning) at 36 tok/s on 2× NVIDIA DGX Spark — W4A16 GPTQ + NVFP4 KV + EAGLE-3 speculative decoding on vLLM. Three serving lanes: speed / balanced / long-context.
-
MiMo-V2.5-TP2-1M-NVFP4-KV-2xDGX-Spark
MiMo-V2.5-TP2-1M-NVFP4-KV-2xDGX-Spark PublicMiMo-V2.5 Omni TP=2 on 2x DGX Spark · 1M context · NVFP4 4-bit KV (~1.97M-token KV pool @ 1M, ~30 tok/s) · 69-eval: thinking-OFF 97.8 beats thinking-ON 90.6 for tool/agent work
-
deepseek-v4-flash-2x-spark-1m
deepseek-v4-flash-2x-spark-1m PublicDeepSeek V4 Flash @ 1M token context on 2x NVIDIA DGX Spark — production-tested recipe (45 tok/s decode, real 800K prompts served)
If the problem persists, check the GitHub status page or contact support.


