[ RETURN_TO_TOPOLOGY ]
user@portfolio:~/blog
[ BACK_TO_ARCHIVE ]

Qwen3.8:27b, a new era of local workflows?

qwen3.8:27b, localai, ai, engineering, thoughts

Introduction

Hi,

Interesting times we live in! I am already a couple of days late testing this thing out, but at first glance Qwen 3.8:27b really seems to be icing on the cake that Qwen 3.6 family already brought us.

General details

According to the Model Card, the scores have improved quite generously across the board. Now you should always take these with a grain of salt, because you really cannot evaluate a models performance purely on benchmark tests. This is due to the developers training the models heavily to get higher scoring, but these values are usually good enough to get some general idea.

  • Text Performance has gone up 11.09 points on average (9.7 median) in Coding, Agent and General combined.
  • Visual Language Performance also improved by 12.21 points on average (12.25 median).
  • Adjustable reasoning effort levels

You can check technical details from the Model Card if you are interested.

Testing the model on RTX 4090

I decided to run this with Ollama this time around as I would have to update my custom fork of llamacpp with context rotation quite heavily to get it working with this.

After downloading the model with ollama pull qwen3.8:27b, obviously here we want to make sure the model loads 100% to VRAM, so we set num_gpu 999. Also num_ctx 65536, since I couldn't quite fit 131k context to my VRAM with gdm3 running (kv cache is q8). I didnt yet test running headless but will update this post once I know for sure the biggest context size you can comfortably fit in. You can find the customizations (Modelfile and kv cache specific changes) at the end of this post. My testing found around 115k context to be the sweet spot running headless.

Cline

So after creating our customized model, I plugged it first into VS Code extension called Cline. It has been my go-to for basic local only agentic workflows due to being quite easy to setup, runs in VS Code, is open-source and is honestly pretty nice piece of engineering.

First thing I gave it was some mundane task of building a website. After getting the plan down, I let the model take over.

It became quickly apparent that the model likes to overthink, a lot. It went into these long thought-processes before doing anything, but when it eventually did, the things it generated were phenomenal.

After our task was done (creating a complete website with very specific details), I noticed an issue. It was not an issue like something was broken. No, everything worked perfectly and was exactly as requested. It was an issue in my own wording and after writing a follow-up specification, off it went again and in couple minutes this website was just as requested down to the miniscule details that would have gone unnoticed in older generations.

So this type of testing is pretty basic and much of it can be pretty easily imitated with iterating on smaller models, but I feel like the amount of details we got here in 2 steps is Opus 4.6 -tier. Which is a model I have used quite extensively, as it is the best model from Anthropic that Antigravity IDE offers for the Pro -tier. (got a free year to use it.)

br.ai.n

Moving on, to gather bit more performance information about this model, I had to update br.ai.n (my own conversation based automated pipeline to build software, I will write a post about it later) and after implementing a few patches and new features to the system, we are ready to run.

So to get Bob (the orchestrator of br.ai.n) working, our workflow is simple. Have a conversation about what you want to create or add to an existing system, after that you give Bob a command to start the factory pipeline and off it goes to generate a plan and build it. Again, this needs a post of its own to explain it thoroughly.

I started explaining an idea I had in mind to build and here I analyzed the output and noticed a quite different conversational vibe to older models, it is very precise to follow the instructions it is given, so the system prompt I have given to Bob, it is following that perfectly. This also seems familiar with my earlier point of the model being very precise about wording in Cline.

So after debating with it for a bit about the implementation, I triggered the build pipeline and watched it work closely from logs. It is a very insecure model, but it is a good thing as it seems to triple check things which then catches issues it might have introduced (which is likely due to kv cache being q8 and the model itself as q4, these will cause inaccuracies). My pipeline has already worked with previous generations of models, generating working software automagically and it is the case here as well. The biggest difference I notice is that this one does not really miss its tool calls and cause the pipeline to iterate unnecessarily.

I did run it couple different times for different things and most of the time it finished the build in one iteration, and on my tests always resulted in working software.

Conclusion

In this narrow scope of testing I have done so far, I would say that we already were in an interesting and capable era of Local AI in the previous generation, but we are now living apparently in the era of having anything near Opus 4.6 level of performance running completely offline on your own machine. It is a great thing but makes me wonder if and when we are getting these models under some regulations. As they say with great power comes great responsibility, I am only seeing the great power side of things with none of the responsibilities.

Thank you for visiting, I hope the rest of your day is going well!


Here is the Modelfile and optimizations I run with RTX 4090:

Modelfile:

FROM qwen3.8:27b

# Context and Batch Optimization
PARAMETER num_ctx 65536
PARAMETER num_batch 2048
PARAMETER num_gpu 999

# Reasoning and Sampling Optimization (defaults)
PARAMETER temperature 1.0
PARAMETER top_p 0.95
PARAMETER top_k 20
PARAMETER min_p 0.0
PARAMETER repeat_penalty 1.0
PARAMETER presence_penalty 0.0
PARAMETER draft_num_predict 4

to fit bigger contexts within smaller VRAM space (this causes bit more hallucinations and inaccuracies):

sudo systemctl edit ollama.service
[Service]
Environment="OLLAMA_KV_CACHE_TYPE=q8_0"
Environment="OLLAMA_FLASH_ATTENTION=1"
sudo systemctl daemon-reload
sudo systemctl restart ollama