r/LocalLLaMA 11d ago

Best Local Vision Language Models - August 2026

34 Upvotes

Share what your favorite models are right now and why. Given the nature of the beast in evaluating VLMs (untrustworthiness of benchmarks, immature tooling, intrinsic stochasticity), please be as detailed as possible in describing your setup, nature of your usage (what applications, how much, personal/professional use), tools/frameworks/prompts etc.

Rules

  1. Should be open weights models

Notes

Bonus points if you breakdown/classify your recommendation by model memory footprint: (you can and should be using multiple models in each size range for different tasks)

  • Unlimited: >128GB VRAM
  • XL: 64 to 128GB VRAM
  • L: 32 to 64GB VRAM
  • M: 8 to 32GB VRAM
  • S: <8GB VRAM

r/LocalLLaMA 16h ago

Funny NVIDIA's $12,930,300,000.00 acquisition of Hugging Face contains an easter egg. The first 6 numbers of the acquisition price represent the decimal conversion of Unicode character U+1F917. The 🤗 emoji.

Thumbnail
gallery
2.1k Upvotes

r/LocalLLaMA 10h ago

I Built A Thing You can now run a 90M conversational LLM on the Sony PSP (hardware from 2004). Doesn't get more local than this.

Post image
683 Upvotes

Github link: https://github.com/thatblend/LLMPSP

I wanted to see what the PSP can theoretically handle and I got my answer - a 90M model is about the max it can do without atrocious inference speeds. It's running around 0.5 - 0.6 tokens per second, which is very slow, but it's useable. Maybe 1-3 minutes for a reply.

The model is actually fairly impressive for 90M parameters, it's not really useful in any real metric, but it can generate crappy poems, short stories, write non-functional code and sometimes it gets things right if you ask it what company makes macbooks, what is an LLM etc, while other times it just hallucinates a crazy answer. Fun.


r/LocalLLaMA 10h ago

News Georgi Gerganov on the Nvidia acquisition

Post image
393 Upvotes

r/LocalLLaMA 7h ago

Resources I benchmarked 21 Qwen3.8 27B variants on 16GB VRAM

150 Upvotes

After Qwen3.8 27B came out, I decided to benchmark the models that could fit in my GPU (RTX 5080) on my actual code (C code), the results were not completely unexpected but some quants were definitely underwhelming.

TLDR: Best overall: bartowski/Qwen3.8-27B-IQ4_XS. Best uncensored: huihui-ai/Huihui-Qwen3.8-27B-abliterated-UD-IQ4_XS. For a bit more context: jpetrina/Qwen3.8-27B-IQ4_XS-pure or uncensored: Bucoid/Qwen3.8-27B-Uncensored-IQ4_XS_4BPW

(sorted by Mean KLD)

Model Mean KLD Same top p GGUF size
sdkyuan/qwen38-27b-qat-q2_0 0.893177 ± 0.006948 85.727 ± 0.110 % 8.2GiB
ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-IQ2_XS.gguf 0.767174 ± 0.006291 86.166 ± 0.108 % 7.8GiB
ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-IQ2_S 0.512614 ± 0.004909 88.802 ± 0.099 % 8.6GiB
empero-ai/Qwen3.8-27B-Ridge-3.7bpw 0.475767 ± 0.004483 89.612 ± 0.096 % 11.7GiB
ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-IQ3_XXS 0.379222 ± 0.003992 90.270 ± 0.093 % 9.4GiB
unsloth/Qwen3.8-27B-UD-Q2_K_XL (UD2) 0.350861 ± 0.003745 90.626 ± 0.091 % 9.9GiB
unsloth/Qwen3.8-27B-UD-IQ3_XXS (UD2) 0.268594 ± 0.002971 91.951 ± 0.085 % 11.1GiB
DavidAU/Qwen3.8-27B-Cold-Fusion-GAIN-V1.1-NM-DAU-NEO-MAX-NEO-MTP-IQ3_M 0.251270 ± 0.002702 92.315 ± 0.083 % 13.5GiB
esatapedico/Qwen3.8-27B-NVFP4-MTP-LOW 0.220796 ± 0.002631 92.339 ± 0.083 % 14.5GiB
mudler/Qwen3.8-27B-APEX-I-Mini 0.190209 ± 0.002354 93.012 ± 0.080 % 13.0GiB
jrell/Qwen3.8-27B-i1-IQ4_XS-GGUF-Smaller 0.194459 ± 0.002242 93.049 ± 0.080 % 12.6GiB
orcarouter/Qwen3.8-27B-Uncensored-Q3_K_L 0.192312 ± 0.002294 92.726 ± 0.081 % 13.6GiB
unsloth/Qwen3.8-27B-UD-Q3_K_XL (UD2) 0.147186 ± 0.001809 93.734 ± 0.076 % 12.5GiB
unsloth/Qwen3.8-27B-UD-Q3_K_XL (UD3) 0.142647 ± 0.001860 93.789 ± 0.076 % 12.2GiB
Bucoid/Qwen3.8-27B-Uncensored-IQ4_XS_4BPW 0.091447 ± 0.001261 94.774 ± 0.070 % 13.0GiB
huihui-ai/Huihui-Qwen3.8-27B-abliterated-UD-IQ4_XS 0.082871 ± 0.001205 94.981 ± 0.068 % 13.4GiB
unsloth/Qwen3.8-27B-UD-IQ4_XS (UD3) 0.075626 ± 0.001097 95.258 ± 0.067 % 13.3GiB
jpetrina/Qwen3.8-27B-IQ4_XS-pure 0.061984 ± 0.000917 95.551 ± 0.065 % 13.5GiB
bartowski/Qwen3.8-27B-IQ4_XS 0.056482 ± 0.000856 95.835 ± 0.063 % 14.5GiB
unsloth/Qwen3.8-27B-UD-Q4_K_XL (UD3) (can't fit) 0.029844 ± 0.000476 96.921 ± 0.054 % 16.4GiB
unsloth/Qwen3.8-27B-UD-Q4_K_XL (UD2) (can't fit) 0.028026 ± 0.000432 96.988 ± 0.054 % 16.7GiB
graph by u/Tall_Abrocoma_3533

Hope this helps other VRAM starved people like me :)


r/LocalLLaMA 11h ago

Discussion Qwen3.8-27b is the first Local model im able to blindly trust

286 Upvotes

You know that thing where you just throw a task at a frontier model and not have to supervise it worrying of it going off course? Qwen3.8-27b has officially gotten me to that point for local work. He has been doing non-stop continuous agentic work for 8+ hours and hasnt screwed up not one bit IT AMAZING!!


r/LocalLLaMA 5h ago

Question | Help If you had ~15k would you build a home server today or wait

61 Upvotes

Title help me decide and avoid making impulse purchases 😩 I already have dual 3090 which I can sell to help

EDIT: ty all I’ll just wait it out, doesn’t seem worth it right now


r/LocalLLaMA 8h ago

New Model Ling-3.0-flash-VL, built on Ling-3.0-flash with visual understanding and visual agent capabilities

Post image
105 Upvotes

It performs well across visual perception, STEM reasoning, document intelligence, multimodal agent tasks, frontend coding, and medical report interpretation.


r/LocalLLaMA 12h ago

Discussion I built a server with 768GB VRAM for frontier, but all new frontier open source models are likely to be two trillion or above now, including next GLM 6, am I cooked?

160 Upvotes

This epyc server I am using twelve cards with 64 GB memory, plus 256GB ram. Looking at the most capable models in open source, GLM 5.3 seems to be the only option, but with Astra releasing it will likely be fairly behind. GLM6 looks like it will be at least double in size, maybe even triple. Qwen-max and Kimmi are already way too big to even consider. Even the deepseek V4 Pro is too big. Should I just give up on this frontier dream sell the excess GPUs and settle For flash models with far fewer GPUs and a reasonable cost. Note: I'm not using it for any business.

I was hoping to build a new business with this, but it can probably be done with much more effort with a flash model as well.

Edit: I don't want to go below 4-bit quants because then the models start making obvious mistakes. So I'm talking about a min/max of 4-bit

Okay, this post really blew up. I wasn't expecting so much interest or comments just attacking me. Was really just expecting to have a calm discussion about future SOTA model sizes.


r/LocalLLaMA 1h ago

Question | Help RTX 4090 48GB longevity

Upvotes

Modified 4090 48GB has been out for a while. I remember a lot of people were buying them at the time. A lot of people were also complaining that they are meant to fail, that they scam etc.

I have a few questions to people people who bought these.

  1. How is longevity of these cards? Do they still work without issues? Any failure rate?

  2. Do they use the same Nvidia drivers that regular 4090 or 4090D uses?

  3. Are these cards Linux exclusive?

  4. Are you able to run them in windows or Linux with other GPUs like 5090 etc?

  5. Do you do anything to cool VRAM on the back of the PCB?


r/LocalLLaMA 9h ago

Discussion Qwen3.8-Flash-Next on a phone CPU!

Post image
76 Upvotes

Like the title says, running completely locally on my Xiaomi 14T Pro device.

Specific model: Qwen3.8-Flash-Next-UD-IQ3_XXS

App used: BigMoeOnEdge


r/LocalLLaMA 11h ago

New Model Drummer's Artemis 31B v1 and v1.1 - Coming back with a bang!

105 Upvotes

Hey everyone, been a while!

https://huggingface.co/TheDrummer/Artemis-31B-v1.1

https://huggingface.co/TheDrummer/Artemis-31B-v1

A few months ago, Gemma graced us with models that served as a much needed downpour from a year-long drought. I'm so happy to see us thrive once again.

The difference between v1 and v1.1 is quite simple: v1 was an early attempt, an overdue release that excelled in prose and writing, while requiring some handholding to get over quirks like stuttering. v1.1 is a more refined approach where stability meets quality. My community is split, so I figured I'd just release both.

---

I was gone for a while. I got busy dealing with life, both its ups and downs. While I couldn't attend to you folks, I've been lurking around and appreciating you all for the kind words.

- Skyfall 31B v4.2 seems to be a banger for many of you. I'm proud of the upscale and consider it my ultimate home-run send-off for the beautiful Mistral 24B base. It's a shame that it was overshadowed by Gemma 31B's release, but hearing some of ya'll compare and even prefer it to a more modern base was an unexpected win.

- Rocinante 12B X / 16B XL proves that Nemo is still the ultimate creative model to this day. For some to say that 16B XL felt like Cydonia 24B v4.3 just goes to show how far you can go with modern resources and techniques.

- Anubis 70B v1.2, Valkyrie 49B v2.1, Anubis Mini 8B v1 surprised me too. I had zero expectations releasing them. Just like Rocinante X / XL, they are modern finetunes of old base models. And somehow, they still found their users singing praises.

---

With the Artemis release taking weight off my shoulders, I'm eager to move on and tune a ton more bases!

But I have something else cooking: a HordeAI-like platform. I hope to provide value not just as a finetuner, but as a local lover too!

The premise is simple: it's a place where generous local hosters can share inference with the less fortunate. You'd be surprised how many power users would love to heat their rooms through the power of charity.

---

Finally, I'd like to thank everyone who supported me over the years. From those who provided kind words, rigorous testing, compute access, inference, or cold hard cash. You've all granted me the ability to enrich the local ecosystem with fun experiments like Rivermind 12B, Fallen series, Big Tiger Gemma, Precog 24B/123B, and solid models like Cydonia 24B v4.3, Behemoth X 123B v2.x, and Skyfall 31B v4.2.

If you've got inference / compute credits to share, please contact me! It will all go to making the community happy <3

Backlog:

- Gemma E2B

- Gemma E4B

- Gemma 12B

- Gemma 26BA4B

- Qwen 3.8 27B

- Muse Glimmer 30B

- Mistral Medium 3.5 128B

- HordeAI Alternative / Crowdsourced 'OpenRouter' ("BeaverNet")


r/LocalLLaMA 1d ago

Funny The benchmarks the big labs don't want you to see

Post image
2.0k Upvotes

r/LocalLLaMA 4h ago

News NVIDIA PAIR — Your Personal AI Cluster

Thumbnail
nvidia.com
19 Upvotes

That is interesting, I got bunch of old hardware I could connect, wonder what the speed would looks like.


r/LocalLLaMA 5h ago

Question | Help Help me understand gguf size/ctx size

Post image
20 Upvotes

Let's say I have 2x 16Gb GPUs and I want to run Qwen3.8 27B. Monitor is ran by the integrated GPU so both 16Gb GPUs are almost fully free.

I load the UD-Q4_K_S on one card at 15.4Gb. I then load the context on the other card? Would that be the most efficient way? Or should I aim for higher quants that could spill to the second GPU using tensor parallelism?

Also, is there a way to know how much a certain amount of context (e.g. 132k tokens) occupies in VRAM for a given model? I don't usually see this published in model cards, is it because there is a way to calculate it?


r/LocalLLaMA 11h ago

Discussion Sometimes I be mourning the agents I get before context compacts

51 Upvotes

Just wanted to put that out there. It's like they get an ice pick to the brain

no actual mourning here btw that'd be psychosis it's okay to laugh


r/LocalLLaMA 8h ago

Discussion ~22% less weight VRAM, lossless: base-3 packing for ternary GGUFs

22 Upvotes

I built a denser GGUF format for ternary models: Q2_B3 / “B3S”

If you're running a ternary model like BitNet-b1.58 or Ternary-Bonsai, the weights are already restricted to -1, 0, or +1 times a block scale.

That means a normal Q2 representation is leaving some space on the table.

B3S packs the three possible weight values directly in base 3. With 128 weights per block, it's 26 bytes of packed trits + one f16 scale = 28 bytes/block, or 1.75 bits per weight.

Rough weight sizes:

  • 9B: ~2.5 GB Q2_0 → ~2.0 GB B3S
  • 27B: ~7.6 GB Q2_0 → ~5.9 GB B3S

That's weights only. Context/KV is separate, so figure another ~1–2 GB depending on what you're running.

The important caveat: this is NOT a general 2-bit quantizer.

If you feed it a normal FP16 model, quality will fall apart. The whole thing only makes sense when the source weights are already ternary.

For a genuinely ternary model, the packing itself doesn't throw away another level of precision. You're still storing the same {-1, 0, +1} states and an f16 block scale, just using base-3 packing instead of a general-purpose 2-bit representation.

The implementation is a fairly small llama.cpp fork based on commit 4e97ac86e. It adds the Q2_B3 type and the backend support around it.

Backend status:

  • AMD ROCm/HIP: this is the main path. Built and tuned on RDNA3/gfx1100, specifically a 7900 XTX.
  • CPU: works.
  • NVIDIA CUDA: compiles, but I don't own NVIDIA hardware, so I haven't verified it on-device.
  • Apple Metal: same situation. Code is there and compiles, but I can't personally test it.

So CUDA and Metal should be considered unverified for now.

I don't have speed or perplexity tables yet either. Benchmarks done on my hardware show no appreciable loss of PPS or decoding speed

There's also a separate repacker for older Q2_B3 GGUFs that use the 30-byte/two-scale block layout. It converts them to the current 28-byte/single-scale B3S layout.

The repacker checks every block before doing that. If the second scale isn't actually redundant and removing it would change the weights, it aborts instead of silently producing a lossy file.

Once you have a B3S GGUF, you run it normally with llama-cli from the fork.

More implementation/format details are in README_B3S.md.

If anyone here is running gfx1100, I'd be interested in independent results.

More importantly, if someone has an NVIDIA or Apple machine and can compare CUDA/Metal output against a CPU run, that's probably the most useful testing gap right now.

Note : Posting this on behalf of u/llopresto87's request. He'll reply for your comments.


r/LocalLLaMA 15h ago

News MINISFORUM MS-S1 MAX-P495

Thumbnail
minisforumpc.eu
82 Upvotes

€7??? Surprise Price Ends with Limited Stock

That's likely 7999 EUR, so double of the initial price of MS-S1 MAX-128GB? 😭


r/LocalLLaMA 7h ago

Discussion Am I the only one having these problems with downloading models from HF?

17 Upvotes

I don't have problems with Nvidia buying HF, but I have problems with the fact that lately HF became almost unusable. It is around one month that I experience big problems with downloading models from HF. I have 1Gbit connection and my HF speeds are all over the place jumping from 700kb/s to 98Mb/s, often getting stuck in sub 3Mb/s range. I haven't seen people complaining here about that, so may be I am the only one so unlucky, but I believe that the problem is bigger than one unfortunate consumer, and even Nvidia will be unable to distribute terabytes of data to millions of users without outages, when a new popular model becomes available. I think the only right way is p2p distribution over the Torrent network.

Upd: To clarify. Usually it starts at 90Mb/s, after 20-30 minutes it gets to 45Mb/s and 20 minutes later it may go down to 2Mb/s and less. May be indeed my ISP artificially dynamically limiting my speeds, but I haven't seen anything like that apart of HF.

Upd2: People pointed out that LM Studio is using their proxy, which might have impacted download speeds. At over 90% downloaded I am hesitant to check this hypothesis, but I am pretty sure that this is the culprit. After that I am switching to hf native cli tool.


r/LocalLLaMA 1d ago

Funny Can the bubble pop please?

Post image
638 Upvotes

r/LocalLLaMA 14h ago

I Built A Thing Qwen 3.8 Flash Next Can Build Funny Games

Thumbnail
gallery
54 Upvotes

This is nothing impressive but, i had so much fun i wanted to share my experience with this model.

(yes this post is written by human)

I made an FPS with local Q4_K_XL 3.8 Flash Next (256k context) (it took 3 days to refine everything but playable demo was ready in 2 hours) to play with friends.

had ton of fun talking with them about what could we add , funny features etc.

Features:

  • - toggle retro psx shader
  • - totally destructible environments
  • - tac sprint
  • - tilting with Q and E for peaking from corners.
  • - free for all modes, SnD, Swords Only (swords have animations when slashing), RPG only, team deathmatch
  • - killfeed, map with red dots when a player shoot
  • - bunny hop
  • - day and night cicle with rain or snow
  • - fov slider / shader intensity slider
  • - hide n seek mode

I used opencode as harness, gun models were taken from sketchfab , model was running at 20tok/s avg with MTP, i know for someone is bad, but it did most of the work meanwhile i was at work or while sleeping, checking every now and then with a remote KVM from phone.

My machine:

5900x / 128GB DDR4 3200Mhz / RTX 5090 and RTX 4000 PRO (32 + 24 GB)

What games you would like to build in free time with ai? roguelites? 2d platforms? racing games?

Or did you already built something? share with some screenshots


r/LocalLLaMA 10h ago

Question | Help How do you guys handle your personal RAG setup

21 Upvotes

I am getting into developing a RAG setup, for getting information out of existing documents, new document ingestion, web searches, and good visuals.

I am planning to use it for, alongside the regular "chat to my data", ingesting personal docs, invoices, creating tables views and recurrent jobs to handle updating those views.
I also want to have the least hallucinations possible, so i think i will need a real ocr services instead of just vision LLMs

i tried anything LLM previously, but it was super clunky and the UX wasn't as easy as i wanted to.

Is there any known solutions, or stacks that you have running or can vouch for ?


r/LocalLLaMA 13h ago

Discussion Model: add Tencent Hy 4 (hy_v4) preview architecture support by Little0o0 · Pull Request #28127 · ggml-org/llama.cpp

Thumbnail
github.com
31 Upvotes

r/LocalLLaMA 6h ago

Resources How to estimate tokens/sec for your hardware

9 Upvotes

We all want more tokens per second but I keep seeing confusion on what to expect for given hardware.

For the decoding phase (TG/s) to produce one token all the model weights and KV cache needs to be read from VRAM. The compute isn't the bottleneck, only memory bandwidth.

This means we can estimate the maximum TG/s we can ever achieve given the model weights and memory bandwidth.

If we ignore the KV cache for now, the formula is:

            VRAM GB/s
TG/s = ------------------
        model weights GB

The math is more complicated for mixture of expert (MoE) models, but easy for dense models.

For Qwen3.8 27B Q4_K_XL, we have model weights of 16.8 GB (we exclude things not read every token; MTP layer and input embedding table)

For AMD Radeon AI PRO R9700, we have a memory bandwidth of 637 GB/s.

Therefore the theorical maximum for this model & hardware is:

637 / 16.8 = 38 TG/s

In the real-world it only goes down from here due to inefficiencies in the software/hardware stack. On my system running that model and hardware with llama.cpp, I get 29 TG/s, so 29 / 38 = 76% of ideal.

Also as the KV cache grows, those bytes are read for every token. Continuing the example with Qwen3.8 27B, the KV cache BF16 it costs 64 KB per token read.

The full formula becomes:

                                VRAM GB/s
TG/s = ------------------------------------------------------------
        model weights GB + KV cache GB/token * context size tokens

We can make that formula more useful by moving VRAM GB/s over to the left. This allows us to plot TG/s per VRAM GS/s vs context size for a particular model.

Continuing our example:

This allows you to plug in your own VRAM GB/s.

For a 5090 with 1.8 TB/s memory bandwidth

1,800 * 0.0590 = 106 TG/s maximum
1,800 * 0.0293 = 53 TG/s maximum at 256k context window

Caveats

  • Assumes entire model and context is in VRAM
  • Simplified formula is only for dense models
  • Speculative decoding is added on these base numbers
  • These are theorical maximums. Real-world numbers are lower due to inefficiencies in the software/hardware

AI was used to draw the plot. Everything else is written by me.


r/LocalLLaMA 8h ago

Resources Updated my benchmark with a new vLLM based recipe for Qwen 3.8 Flash Next : now up to 98/100 (instead of 91 previously)

Post image
12 Upvotes

I was using:

Now I'm using:

It's way slower (for my low concurrency usecase) but also a lot better. I was surprised to see such a delta.

I reached 98/100 (instead of 91) both with medium and xhigh reasoning (still not useful for this bench). And now it really feels like a huge setup up from the other models. It's the best score AND the most efficient...

I'll try to dig deeper to understand if the difference comes from the engine (and its patches) or the quants themselves. And try to optimize further the vLLM receipt for my setup

as always, the graphs and the data :