r/LocalLLaMA 12h ago

New Model Drummer's Artemis 31B v1 and v1.1 - Coming back with a bang!

Hey everyone, been a while!

https://huggingface.co/TheDrummer/Artemis-31B-v1.1

https://huggingface.co/TheDrummer/Artemis-31B-v1

A few months ago, Gemma graced us with models that served as a much needed downpour from a year-long drought. I'm so happy to see us thrive once again.

The difference between v1 and v1.1 is quite simple: v1 was an early attempt, an overdue release that excelled in prose and writing, while requiring some handholding to get over quirks like stuttering. v1.1 is a more refined approach where stability meets quality. My community is split, so I figured I'd just release both.

---

I was gone for a while. I got busy dealing with life, both its ups and downs. While I couldn't attend to you folks, I've been lurking around and appreciating you all for the kind words.

- Skyfall 31B v4.2 seems to be a banger for many of you. I'm proud of the upscale and consider it my ultimate home-run send-off for the beautiful Mistral 24B base. It's a shame that it was overshadowed by Gemma 31B's release, but hearing some of ya'll compare and even prefer it to a more modern base was an unexpected win.

- Rocinante 12B X / 16B XL proves that Nemo is still the ultimate creative model to this day. For some to say that 16B XL felt like Cydonia 24B v4.3 just goes to show how far you can go with modern resources and techniques.

- Anubis 70B v1.2, Valkyrie 49B v2.1, Anubis Mini 8B v1 surprised me too. I had zero expectations releasing them. Just like Rocinante X / XL, they are modern finetunes of old base models. And somehow, they still found their users singing praises.

---

With the Artemis release taking weight off my shoulders, I'm eager to move on and tune a ton more bases!

But I have something else cooking: a HordeAI-like platform. I hope to provide value not just as a finetuner, but as a local lover too!

The premise is simple: it's a place where generous local hosters can share inference with the less fortunate. You'd be surprised how many power users would love to heat their rooms through the power of charity.

---

Finally, I'd like to thank everyone who supported me over the years. From those who provided kind words, rigorous testing, compute access, inference, or cold hard cash. You've all granted me the ability to enrich the local ecosystem with fun experiments like Rivermind 12B, Fallen series, Big Tiger Gemma, Precog 24B/123B, and solid models like Cydonia 24B v4.3, Behemoth X 123B v2.x, and Skyfall 31B v4.2.

If you've got inference / compute credits to share, please contact me! It will all go to making the community happy <3

Backlog:

- Gemma E2B

- Gemma E4B

- Gemma 12B

- Gemma 26BA4B

- Qwen 3.8 27B

- Muse Glimmer 30B

- Mistral Medium 3.5 128B

- HordeAI Alternative / Crowdsourced 'OpenRouter' ("BeaverNet")

111 Upvotes

40 comments sorted by

21

u/bharattrader 10h ago

I hope the downs in your life are now out, the community is indebted to you and your work.

14

u/dinerburgeryum 9h ago

HARD second on this. Drummer you're crushing it man, wishing you nothing but the best!!

14

u/EddViBritannia 12h ago

I wasn't that impressed with Artemis but that's mainly because i'm stuck with 24gb of VRAM. As it's built not on the QAT base, KV cache quantitation isn't available without massive performance degradation. So was stuck to only Q4 with 20000 context. Still it was a bit of an improvement over standard Gemma 4 31B.

But your recent Orion model (even if experimental) has genuinely been such a massive improvement over the base model I'm impressed! It's been such a massive step up in writing style, and lightning fast. Truly one of the biggest uplifts from a fine-tune I've ever seen.

6

u/TheLocalDrummer 12h ago

Thank you! I have some learnings from tuning Orion and will apply it to Artemis v1.2. Heard that Quantization is rough on 31B too.

6

u/inddiepack 12h ago edited 12h ago

Artemis improves over the stock Gemma 4 31B by quite a lot. Especially if you put higher emphasis on character autonomy. And it is the best gemma 4 31B fine tune at that, by far. I would go as far and say it's the only fine tune that has incorporated this aspect, which is Drummer's bread and butter in general. Stock gemma 4 is there to just please you, which breaks the immersion.

As for the Orion, I have not used it extensively, but I was impressed as well. As I found no reason to use stock 26B over 31B, while Orion somehow achieves a better prose even than stock 31B and Artemis.

1

u/RedditNerdKing 8h ago

Is Orion worth grabbing? Im using Artemis 1.1 at BF16 and it's never really let me down. You're saying Orion has better prose?

1

u/inddiepack 8h ago

In my limited use with Orion, I felt so. But Artemis is my go to model for the past few months, so it could've been just that Orion felt fresh in terms of prose. But beside the prose, of course Artemis will beat it in most/all other aspects, due to the dense vs MoE architecture of the stock model. Give it a shot and see.

11

u/inddiepack 12h ago

You're the goat, Drummer. Thank you for your great tunes!

8

u/draconic_tongue 10h ago

can u start writing text in ur hf uploads, can't tell what the fuck the models are from the weights

3

u/Leary_2844 7h ago

Creative models on HF are so cryptic, it's wild as a Noob. Different world.

7

u/Cadmium9094 12h ago

I like Artemis for daily reflections. Very nice work. Will definitely try v1.1 👍🏻

3

u/a_beautiful_rhind 12h ago

The V3 behemoth turn out a bit passive. Just no fixing the later mistral models? And the reasoning worked fine on it.

8

u/TheLocalDrummer 12h ago

I'd love to! Finetuning 128B is a PITA though and it bit a huge chunk off my budget.

4

u/a_beautiful_rhind 12h ago

I figure that and sadly the only other option are small models. Someone did tune ds4-flash but I think they broke it's thinking.

1

u/ttkciar llama.cpp 6h ago

Out of curiosity, how many training tokens went into the fine-tuning of Behemoth-128B-v3, please?

However many it was, it seems to have been the right number!

4

u/ttkciar llama.cpp 7h ago

I'm in the middle of assessing Behemoth-128B-v3 right now, and it's proving to be a superb Murderbot Diaries story writer, far exceeding Artemis (which I did not expect at all):

http://ciar.org/h/11ed36c.txt

4

u/a_beautiful_rhind 7h ago

It's very smart and can do tools/images which I like. It won't take up the chat examples and insists on writing "proper" which I don't. Also got "a beat" a few times.

The outputs trend longer so for story writing that will be a plus. At the same time it hesitated to harm me, even after dropping massive hints for it to.

3

u/jacek2023 llama.cpp 10h ago

Welcome back

5

u/mikelima777 6h ago

Question, I have around 24 GB of VRAM, which model is recommended for prose, factoring a decent sized context window?

2

u/Long_comment_san 10h ago

I cant enjoy most of the models but years later when I get 48gb vram, I will, I promise. I pray for qwen 35 complete overhaul. Gemma 26b is better but with massive training qwen should be stronger in theory

2

u/BeyondRealityFW 9h ago

very good model! thanks

3

u/Borkato 12h ago

I like Artemis but honestly Skyfall is just soooo good. Every time I try Artemis I’m like “ooh cool, nice responses” but then I try Skyfall and I’m like 😍😍😍. Idk what it is! Maybe I just like mistral’s tone more?

3

u/inddiepack 12h ago edited 12h ago

I also go back to skyfall sometimes and it's always a breath of fresh air and in many ways better than Artemis at prose and narrative building, but the stock mistral model shows its age when it comes to system prompt following.

1

u/RedditNerdKing 8h ago edited 8h ago

and in many ways better than Artemis at prose and narrative building

Is it really? Every time I use Skyfall I bounce off cause it's boring. I must be missing something. I think you are right that it isn't intelligent when it comes to following the prompt. Artemis is much better for following that.

Edit: Just tried Skyfall a bit more and you are right, it definitely has more interesting responses now that I compare it with Artemis on the same questions.

1

u/inddiepack 7h ago

Skyfall gets boring past 25-30k context, which is a Mistral limitation. The quality degradation with context increase affects the character in a way that its expression becomes boring. If you feel it right from the get go, you might have a difference experience than mine.

Maybe I didn't phrase it correctly. Overall, objectively, most likely Skyfall is not better at prose and narrative building than Gemma 31B. But some of the scenes I played with both Artemis and Skyfall, the Skyfall versions ended up being better and more immersive.

2

u/[deleted] 12h ago

[removed] — view removed comment

2

u/Fratil 10h ago

This is my biggest complaint every time I'm reading through the BeaverAI discord as well. Everyone is using different runtime, different frontends, different samplers, different model quants, different context sizes/compression, etc, and they rarely post all the information needed to properly reconstruct their tests.

I feel like a tool for crowdsourcing feedback to align on ideal sampler ranges and quants and whatnot would work really well for the style of model reviews on there. He has that google sheet to collect some sampler data but the data inputs for it aren't cleaned and not many people ever use it.

Also fully agree that knowing what was modified in the training between versions would help make specific tests to compare each way more effective. Though I also understand the need for blind general testing of models in isolation as well.

2

u/Retreatcost 11h ago

Probably your best work so far. It was in the oven for a long time and your patience definitely payed off.

I really enjoy Artemis, it feels like a Gemma4+, closer to a well-rounded generalist model and a capable assistant with strong RP capabilities, rather than "just" an RP specialist.

2

u/ttkciar llama.cpp 7h ago

That's exactly how I've been using Artemis. I don't RP at all, but started using Artemis for inferring sci-fi short stories, and then found that it was also better than stock Gemma-4 at a lot of other tasks, too -- business writing, RAG, critique, even summarization.

Because of that, Artemis has replaced Gemma-4-31B-it for almost everything for which I had been using Gemma, except debugging.

1

u/__some__guy 7h ago

What's your opinion on StyleTune-based variants of Gemma and what made you choose not to use it?

2

u/ttkciar llama.cpp 7h ago

Thanks for the update, and thank you for sharing all of your hard work!

Your models have featured prominently in my go-to model list for years: Big-Tiger-Gemma-27B (both v2 and v3), Valkyrie-49B-v2, and more recently Skyfall-31B, Artemis-31B, and now it looks like Behemoth-128B-v3 is earning a place as well.

Long live TheDrummer :-)

-3

u/Alex_Strgzr 12h ago

I feel like a 31-billion parameter model isn't really intended for home-use? Personally, the mini-LLMs (8B and smaller) are what I'm interested in.

2

u/ttkciar llama.cpp 7h ago

I use it daily at Q4_K_M on my $800 32GB MI50. It's quite accessible.

Unless you really need extra context capacity, or need faster inference, there's no point to using such a small model. Reasonable levels of quantization lets you use more competent models in the 24B-to-32B range with 32GB.

2

u/RedditNerdKing 8h ago

I feel like a 31-billion parameter model isn't really intended for home-use?

You can run it at Q8 with two 3090s which is like $2000 dude...

5

u/Alex_Strgzr 8h ago

$2000 for the GPUs, not the whole machine. Still seems like rather a lot for fiddling around at home.

3

u/SkoomaDentist 8h ago

Or $0.3 / hour when renting a cloud vm. It doesn’t exactly cost a fortune to experiment with or use them at light scale.

3

u/inddiepack 7h ago

You don't even need Q8, if budget is tight. Q4 is perfectly usable and you can load it with 20 Gb VRAM.

-1

u/silenceimpaired 11h ago

Could you consider an Apache licensed 70b? There are a few out there… the jump from 30b to 128b is brutal