Bonsai 27B: A 27B-Class model that runs on a phone

7 days ago (prismml.com)

What I most want to see it compared to is Gemma 4 12B in the 4-bit QAT version. It's barely bigger than this at just under 7GB, so it also runs on just about any modern device and is remarkably smart for its size. It's an excellent tool user, crazy good vision for its size. I'm still trying to wrap my head around how much is lost with each step down in resolution, but the QAT versions from Google seem to prove the answer is "very little" at four bits.

  • Based on their numbers and cross referencing with the Gemma numbers, this model crushes Gemma 4 12b on math and coding, is slightly worse on knowledge and tool calling, and is significantly worse on vision tasks.

    • I think this is where leveraging classifier models will become important. The frontier LLM models do "everything", while we've known for a while that to truly scale this we will need to distill models into their individual functions. I don't see this as necessarily a bad thing and hope more is done in this space. Very promising.

      12 replies →

    • FWIW my tests on my little puzzles suggest that it is not better than the Gemma 4 12B on SQL. It really does seem to get quite tangled up on stuff.

      PHP/Wordpress code seems OK (better than the Gemma) but it gets stuck in reasoning loops.

      Mind you, I am something of a cynic about the underlying 27B dense Qwen; I think the 35B MoE model is often better and it is just so, so much faster.

      2 replies →

    • From my own experiments with local, low VRAM model use vs. what I'm used to from using Claude at work is that being good at "coding" is of no use, if you're worse at "tool calling" as coding in an agentic way requires quite a bit of tool calling.

      If you can "hide" different models of 8GB VRAM requirements each that have those specialties and mix and match them for me without having to manage it manually, I'll be impressed. Until then I will keep using my Claude, because "remarkably good _for their size_" models I've tried so far just sucked at trying to use them the way I code at work with Claude.

    • > slightly worse on knowledge and tool calling

      Worse than Gemma at tool calling? Gemma's already bottom tier at that (at least when there's Qwen to compare to), that would just be unable to do tool calling at all.

      2 replies →

    • To be fair, everything (roughly within an order of magnitude in size) is worse on vision. 12b is a beast for vision tasks, better than its bigger siblings, even.

    • More to the argument that we need a model of models - one general one that calls specialists in to do what they are good at and handles that like a foreman for you.

      4 replies →

    • The things it loses are all the things that google models are historically excellent at, so that's a reasonable performance. I think the take home here is that the 1 bit models are probably better, but it's not a slam dunk given advanced quantization techniques.

    • The Gemma models are so good at vision. It seems particularly important for phones. Also, they write in a much more pleasant manner than Qwen imo.

      1 reply →

  • 4bits is a cutoff point for many model families, but also depends on what parts you quant to 4bits vs alternatives (weights, weight+activation, kv cache). Also depends on model size and task, lots of nuance in quanting I've come to learn.

    Good evaluation from 2024 https://arxiv.org/pdf/2402.18158

    I'm currently working towards an updated version (not an og author), curious if others are aware of similar surveys, as I have yet to do a real lit search.

    • The key point here, I think, is not the 4-bit but the QAT — the model is trained with the objective of losing the least at 4-bit quantiZation (I am assuming it is literally about assigning numbers that quantize better).

      The 12B QAT model is indeed sort of mindblowing.

      7 replies →

  • I'd like to see them do a 1-bit binary version of Gemma 4 12B ;)

    • I want 31B. The 12B 4-bit QAT is already small enough to run well enough on every device I use regularly, including phone and tablet, I don't need a 1-bit or ternary version of the 12B.

      But, what I really want is for Google to release bigger Gemma 4 models, particularly a bigger MoE, like a ~70B or ~120B. Gemma 4 is the best all-rounder among the models I can self-host even though I've got a 128GB Strix Halo. A 4-bit QAT version of a 70B MoE would probably be the sweet spot.

      A bigger Qwen 3.6 with a 4-bit QAT version would also be welcome, as the prior bigger versions aren't notably better than 3.6 27B, but I guess Qwen is done doing larger open weights models. They did release AgentWorld recently, a post-train of the 3.6 MoE, so they're still doing some open things.

      I think I want to see more third-party testing of this ternary Qwen to know if crushing it to 1.56 bits kills it; there are tons of benchmarks of Qwen 3.6 27B, so it's an ideal candidate to figure out what the extreme compression does to it.

I need help understanding this. I understood that the magic here is the quantization that allows it to use from 50G to 4G and their process retain most of the intelligence within Pareto limits of gain. And then they proceed to compare with other quantized models as in the level of intelligence per size. It gets to my attention though that the performance in tool calling is mostly affected which is a problem for other small models.

How does this model compare to a recent 4G model? How do we know it retained intelligence from the parent rather then being fine tuned for the benchmarks?

I am not shtng on them or anything. I'd rather find it amazing, BUT given my limited knowledge, I feel the results miss fair comparison plots and the ones might be misleading. Buy I also reckon it might be me the problem. Anyone care to explain this poor silly fellow some of those points?

  • from what I understand prismml isn’t doing a quant like normal models where you take a model trained at fp16 and then chop off some bits to reduce vram, but rather they’re training the model natively with 1 bit weights. It’s explained more in the article. They’re also doing some other tricks like a fp16 weight per block of 128 1bit weights to get some more data out of 1 bit weights

    • They aren’t training at all. They are quantizing existing models, it’s just that the process is different. The 27B uses Qwen3.6 27B as the base model.

  • 1-bit weights are not pointers so cpus can process them, storing them takes less space etc. tons of gains

Apparently Apple is "in talks" with the PrismML: https://www.cnbc.com/2026/07/14/apple-prismml-ai-compression...

The models themselves are showing up on Hugging Face here: https://huggingface.co/prism-ml/models

I've tried a couple in LM Studio - the GGUF one and the MLX one - but neither worked there. Anyone else get them to work? Might be that LM Studio needs to upgrade their llama.cpp or MLX engines first.

  • I can report that it's working in oMLX. I've been experimenting with the ternary one; it is quite an impressive model! I've been grilling it on some deep learning/computer vision stuff and it's aced everything so far. Responses are thorough, accurate, sophisticated. General knowledge outside of CS doesn't seem as robust, which I expected. Honestly, I don't think the examples in the blog post do it justice.

    • Nice, I tried it too with oMLX — agreed, it seems very capable for coding! Was slightly underwhelmed by performance though. I got about ~24 t/s on the ternary version on my M2 Max 64. That’s quite a bit slower than Qwen A3B 35B (4 bit unsloth). How was perf for you?

      5 replies →

  • I got their previous model working in their custom fork of llama.cpp (https://github.com/PrismML-Eng/llama.cpp). I haven't tried this one yet, but will find some time to benchmark it sometime this week.

    Though this says mainline llama.cpp has their patches for Metal and CPU backends, so maybe it's simply "use current llama.cpp" if you have a Mac or fast enough CPU/memor to use the CPU backend.

  • Not sure if there's any way to run Prism's fork of llama.cpp inside LM Studio.

    The fork runs fine for me. The model gets very notably stuck in a reasoning loop on one of my simple tests, though it might be that it has the same issues with setting reasoning effort high.

    On my M1 Max I still think the MoE Qwen 3.6 and Gemma 4 models are the best options. And I am far from convinced that the 35B is actually worse; it gets stuck in reasoning loops much less often than 27B in my experience.

I have benchmarked Bonsai 27B CPU inference on my computer (a Ryzen 7 5700X desktop with 48G RAM running Ubuntu 24.04) using the latest 62061f910 build of PrismML's llama.cpp fork.

Binary: 9 t/s prompt, 6 t/s generation. Ternary: 0.8 t/s prompt, 0.7 t/s generation. It looks like CPU inference for ternary isn't optimized yet.

Excuse my likely stupid question, but has anybody had some success using Claude Code with frontier agents (or Junie or anything else) to invoke local LLMs for specific sub-tasks or wrapped as skills? In other words, is there a way to use expensive, frontier models as orchestrators that manage local models to do the specialised coding tasks?

  • That is my entire workflow in opencode through delegates. I have an "orchestrator" which uses matt pocock skills like grilling and writes specs in gh tickets. Once that is done it delegates individual tickets to deepseek v4 flash based developer subagent which executes the code pretty fast. This agent can only delegate one subagent which is a code-reviewer which uses a better model like qwen 3.7 for review. So dev does its own review loop before returning.

    It's going amazingly. The orchestrator holds the big knowledge from grilling and also enables me to do more grilling to refine the specs. The trick was to check and re-align after each phase was developed. Also carefully defining each workflow step explicitly otherwise it makes mistakes like trying to self review the code. Also needed to define very explicit contracts. I chose the same surfaces as the matt skill set.

  • I'm working on the area in https://beolis.com.

    The system as a whole is meant to support that use case, where each task (ticket in its jargon) can be tackled using a custom workflow that can each use a different agent/llm (so, it should support local LLMs if you have configured your coding agent to use them).

    Sidenote: it's still not where I want, but getting there...

  • This may not be exactly what you're asking, but I've been running Hermes using different frontier models (like DeepseekV4-Pro, GLM-5.2, and others) as the heavy lifters but who also spawn agents to run my local models (mainly Qwen3.6-IQ4) for offloading tasks that are a good fit for a quantized model that's local. Works really well.

The KV-cache memory usage also seems remarkably frugal, even at the full context length. That could make this model particularly useful in multi-agent coding workflows.

I wish KV-cache memory usage and related optimizations were discussed more clearly in new model announcements and demos.

  • quanting kv cache hurts attention / recall, and long-form tasks by proxy. Model families and sizes have different tolerances to quant ting different parts of the model, same for intended tasks.

I find it super interesting that we're now in an era where we have LLM's that are quantized to binary weights - 1's and 0's. So effectively they're digital neural networks.

I assume that in addition to the significant memory savings, this should also lead to much simpler matrix multiplication operations? Could models like these run on CPU's efficiently, or does the geometry of the compute mean GPU's are still a better choice?

  • Its not about geometry, it's a parallel compute thing thing. CPUs typically don't have more than 10 or 20 cores. GPU have 100s to 1000s.

    Matmul is very well parallelized. More, lower power cores will always pay off handsomely.

  • Most of the time, the speed of these models are constrained by memory bandwidth. GPUs normally have much more memory bandwidth.

    • I'd expect the memory bandwidth to be the same for the CPU and GPU under a unified memory architecture like Apple silicon uses?

TIL that 1 bit models are actually 1.58 bit with three values +1, 0 and -1

  • There's two variants of this (or, as the joke goes, for very big values of bit):

    Ternary Bonsai 27B uses ternary {−1, 0, +1} weights with FP16 group-wise scaling, giving a true 1.71 effective bits per weight.

    1-bit Bonsai 27B uses binary {−1, +1} weights with the same group-wise scaling, giving 1.125 effective bits per weight.

    • this is a really dumb question, but how is -1 represented?

      is it a float? if so, how many bits is the float?

      I've never heard of a bit ever having more than two possible values

      7 replies →

  • 1.6 bits if you want the most practical way to pack five 3-state numbers into a single byte. But even then, they usually pack four 4-state numbers instead.

  • Yeah, it's an unfortunate convention from the very first "1 bit" model. But to be clear, Bonsai comes in both ternary and actual 1-bit variants.

maybe its nitpicking here but the demo shows them asking the model what to cook and its recipie sounds like it wouldn't be very good and also it totally gets the macronutrients wrong. 25g protein for "spaghetti, carrots, peppers, garlic and herbs"?

  • And why would i want such mundane questions to be handled by an AI on my phone? That sort of thing doesnt need AI, let alone a local one. Basic google search was answering those questions long ago. My point: phone-sized AI is only useful if it can do things that only AI can do. Can it ingest a document scanned by the phones camera? Can it translate in real time? I dont see how or why i would ever ask it for recipe advice. That need is met elsewhere x10.

    • I had exactly this use case - in a grocery store in the Alps, no internet, fired up a local LLM on my phone to figure out what to cook and what to buy

    • Awwwe, cmon. You are thinking about the problem space incorrectly. This is an opportunity to create a unicorn company that develops the first AI tongue.

  • I personally don't like carrots much, but it doesn't sound bad to me – could definitely use a tomato sauce but that doesn't seem to be an option from the image it was given.

    > 25g protein for "spaghetti, carrots, peppers, garlic and herbs"?

    Maybe it assumed the pasta was some kind of protein chickpea pasta? =P it definitely seems wrong.

    • Many people don't realize that regular spaghetti actually is pretty high in protein. All that gluten.

I've been watching and waiting for this, interested to see how smart it is, as it fits with my interest of getting the smartest possible model running in 10GB of VRAM (RTX3060 that has to drive 2 monitors and run an llm)

  • Toss the rtx into a cheapo optiplex or thinkcenter, and run it headless - the load on your machine is gonna make doing other stuff while it’s running painful. Plus that frees up the rest of your vram.

    • Running an llm on a PC running a desktop environment really isn't that bad. You lose a bit of ram & vram but that only matters if you _reaaaally_ want to push to the max model size your hardware can handle.

      The biggest issue I've found is absent mindedly opening YouTube or the like that spike ram requirements and freezing the system up. But that's a me problem

    • I aspire to someday move up to an AM5 system, but for now, the Dell T3610 has to do everything. Might get a second GPU for it though.

The problem, of course, is if you run the UD_Q2 variant (Unsloth) which does only post-training, the number is pretty close to 1-bit model here and the 5% drop in tool-call is significant than it suggests in real-life use cases.

  • You also need to pay close attention to BFCLv3 multi-turn result, that helps you to get a sense how frequently these quants will be in a doom loop.

Got this running on my phone. Unfortunately like other small models it hallucinates quite easily.

eg asked it what Signoz is. It reckoned it is a woocommerce/shopify competitor aimed at India market

Tried it on Android and got "!!!!!!!!!!!!!" for answers.

  • The qwen models really seem to have this as a failure mode, its so annoying having a proper trace ending up in !!!!!! Garbage.

    • Oh, interesting. "!" is token id 0 in the Qwen tokenizer; I wonder if there's some tokenizer shenanigans either in inference or training that end up causing this specific behavior.

    • Wait in a regular sentence, what is the probability of "!!!" being followed by "!"?

      Sounds like the model is not following a proper probabilistic choice here, so maybe more a programming error than a model training error.

      1 reply →

  • That's what happens when you quant too hard. I'm working on quant strats and evals for the same underlying qwen 27b models.

    When I saw 27b on a phone, I thought not fitting, big phone, or aggressive quant. NVFP4 still takes 27G before KV cache.

For those curious about their demo, I’m pretty sure it’s using Locally AI (iOS only) that lmstudio acquired/aquihired a couple months ago.

After using a highly capable 2-bit quant as my daily driver for months now, I get pretty excited about releases like this. After a few days for the kinks to be worked out, I’ll be excited to try it.

  • What model? And what hardware do you run it on?

    I find these style of models are great, but fail hard, and fail randomly. I'd be hesitant to use it for a daily driver, but I'm using dual 3060s, so it's not like I'm quantizing a frontier model here.

    How do you find the overall experience? And do you have any special sauce or recommendations for going this route?

    • I’m using DeepSeek V4 Flash on 128gb mbp - it’s a bit different using a 200b+ param model. It’s MoE so performance is acceptable. It will still malform a tool call every now and then, but the capabilities are so far ahead anything else that the majority of the time it works really well and solves really complex problems.

    • I've got dual 3060s as well. What's the best models you've found for this setup?

This is accelerant #3 and #4 from our article converging in one release: a 27B-class model, built on Qwen (already one of our examples of local models "good enough to matter"), now running on an iPhone. The hardware layer and the local-model layer aren't just going to converge in the future, they're doing it right now! https://news.ycombinator.com/item?id=48892559

Preliminary analysis via lm-evaluation-harness + vllm

    model         | disk | wikitext | gsm8k (match/error)
    baseline      | 55G  | 8.00     | 0.50/0.09
    nvfp4-gptq    | 27G  | 8.25     | 0.47/0.9
    nvfp4a16-gptq | 27G  | 8.11     | 0.53/0.9
    bonsai-4bit   | 19G  | 16.75    | 0/0 (eval bug?)

Looks like they quant'd too hard at 4 bits, can't imagine the ternary being any good based on this. I'm also not sure what is up with the gsm8k, their benchmarks show something different, but they are using another eval tool. I'll have to add it to my setup. Also why I'm building a setup instead of taking model devs word for benchmarks. (https://github.com/modelscope/evalscope)

Code if you'd like to reproduce or try other test sets: https://github.com/verdverm/quantr (lightly tuned to a single oem spark, probably possible in 32-48G)

Good paper to understand the effects of quant regimes across model families and tasks: https://arxiv.org/abs/2402.18158 (Evaluating Quantized Large Language Models - 2024 ICML)

  • Doesn’t this suggest you aren’t properly running the model?

    • Its unclear if its in vllm, the eval harness, or the weights. I'm running it the same as other qwen3.6 derived models and it appears to work, but other comments speak of waiting for support in their tool. Could be a chat template or could be legit because I'm using a different quant / bit selection from their hugging face and using their primary model as the comparison point for my huh? They may have put more effort into the flagship model than the 4bit because they are focused on on-device model running.

      Software was already complex, now we are adding highly non deterministic elements into the mix

I still don’t see the point of this. In my testing, it’s worse than Qwen 3.5 4B and even 0.8B.

  • When new models are released (I realized this is qwen 3.6 but the quant is novel) - it takes a few days for the kinks to get worked out - you’ll likely have better luck if you give it a few days and try again.

  • It is still worth experimenting. Although we do need independent evaluation of these models instead of labs posting such biased and skewed results.

What's the hiring space and business strategy around all of these smaller AI labs? Its really cool that people like these guys get paid to optimize models and give them out for free (open source). Do a lot of these labs have forward deployed engineers doing integrations with customers who want local models? Is there a general shift towards the local model crowd?

  • If you read to the bottom of the page, it says they're funded by a few people, and one of them is Samsung. I'm betting Samsung wants to be able to ship a capable AI system on a future model of their phone so they can compete with Apple.

    • Agreed, and the prevailing wisdom now seems to be that unless you can release a truly frontier model, you might as well release yours as open source to undercut your competition.

Quite weird that heavy quantization method on a dense model gives better results than slightly quantized MoE models like 35B-A3B from Google.

At this point all the different quantization and 'compression' (look at MPO applied to LLMs...) techniques start feeling a bit like snake oil. It's just gut feeling - or scores on benchmarks models are optimized for - what ends up deciding whether a technique is good enough or not.

So first off, phenomenal stuff to see a 1bit model at 90% capability.

However, this is the 5th product post in 2 weeks that proclaims that AI use is shifting, and why [insert tradeoffs] are the perfect fit.

Paradigms shift don't happen in the release announcements.

I suspect this is an AI-ism, making all the release posts sound so paradigmshiftery.

  • That is true of all tech announcements. Marketers will do their thing regardless of reality.

I don’t know if the llama cpp implementation is wonky (and only supports the binary version) but it’s a lot slower than 35B-A3B @ Q4_KM + MTP with CPU offloading.

  • It’s still a dense model so all 27B parameters are active. 35B-A3B activates only 3B at any time.

Tried this on an old 4 core i5 and got about 1tps.

OS: WSL2 on Windows 10

What is the best way to deploy these on CPUs, arm64 ones in particular?

I'm interested in the CPU inference application of these models with things like the FairyFuse kernels.

I've tried trillim previously but was disappointed that i got higher tok/s just with similar sized models through ollama using just Q4_K_M quants.

I see there is bitnet.cpp and litespark-inference. What else should i look at?

  • Llama.cpp should be rolling out support for this soon if they haven't already. Cactus is a but more targeted for efficient ARM execution, but I haven't been keeping up with what all they support. Would be worth a feature request if they aren't working on it yet.

Entire blog post seems to be AI-generated :/

  • Do you think people who work on AI for a living are not going to use it?

    • Of course not, personally almost all of my code these days is generated.

      The LLM style of writing is just very distracting to read. “It unlocks X”, “Y changes the equation”, and why is there always something shifting? Makes my eyes glaze over in an otherwise interesting post.

      1 reply →

    • IDK, Anthropic's public posts are mostly free of Claude-isms. (Their documentation less so...)

How does one even install and use this on a phone? They don't say. I guess the claim is fake. It is unreasonable to make a claim without an app for Android and iPhone that supports and runs this model.

I tried this on M1 Pro today with 16GB ram and it worked!!!

I was using vscode and it seemed to interperet the system prompt right and then started actually inspecting and doing stuff.

Unfortunately the vscode system prompt is 24000 tokens, and I was getting 100 at beginning, 69 by the end of it, but honestly I'm super impressed. Great work team 1

27B is way more than you need for a phone. Doesn't matter how much you try to compress it, it's the wrong application of the wrong tool. There are already useful tiny models that fit on phones and do basic things really well. Dumb down a big model too much and it becomes worse than a small fine-tuned model.

Thanks Bonsai team.

Now open weight LLMs/VLMs/LMMs are becoming even larger to the extent that consumer-grade hardware are no longer able to run these models. In contrast, quantization and pruning make the model better at the size-performance pareto and provide people with strictly more possibilities.

Nice to see a larger model in their lineup, I've been using Ternary 8B and it seems to get higher TPS than most other similarly sized models on my hardware.

Does anyone know how to disable the 10 minute timeout in Android studio when using a self hosted model on local machine?

It always auto disconnects.

From an investors perspective, this is truly a paradigm shift - this will kill a whole range of startups in Europe which were packaging privacy and wrapping around large hosted models. There's absolutely no reason to use a "Privacy GPT tm" provider, then I have it all on my own laptop - There is also no need for banks or other regulated institutions to rely on those providers when they can selfhost with this much intelligence on tap.

  • It will when they can get the performance up a bit.

    My brief experiments with the ternary version suggest that it broadly meets their claim to be a 27B model that fits in much less RAM, that is for sure. It is about as fast as the underlying Qwen 27B but it gets stuck in reasoning loops quite easily.

That's awesome. What's the largest model that could fit onto a single 16gb gpu at 1.125 effects bits per weight?

  • Doing some naive math, the F16 filesize is ~53.8gb, the 1-bit version is ~3.8gb, about 7% of the original size. The F16 size is roughly 2x param count, so that gives a rough ballpark of ~110B.

    • Which would be very interesting to test, as larger models (such as Deepseek V4 Flash or Qwen 397B) seem to compress better. Their Q2 quants are usable as is, even without the ternary compression.

  • Yep, that’s the question. I asked just that when Bonsai’s first models got released. Super interesting if we can push the parameter count over 100B with 1.125 bit quantization and still keep pretty good performance versus 16-bit 100B models. That’s a definite sweet spot.

Nice!

Do they have plans to bring even bigger models down to ~16GB VRAM so that more consumer hardware might be useful?

  • bigger quant'd harder is not always better than a model of more modest size and quant

What data type stores one bit? Does it offer opportunity for more efficient matmuls?

The meal demo is hilarious.

— Hey, model, see this fake-ass stock photo of a variety of spices, vegetables, and spaghetti? What meal can I make with this?

— Just cook everything.

— I’m a complete noob. I can’t even fathom how to cook those things. Help me!

— Sure sure. First boil the spaghetti completely and drain. Only after that, while it’s getting cold, you need to sauté (good luck knowing what that is if you don’t even know how to cook spaghetti) the garlic and carrots at this specific temperature (good luck figuring out how to do that on a stove). Despite having mentioned the peppers and herbs in the previous message, I’m not going to tell you what to do with those. Just chew them raw or something, I guess.

The demo shows that the model can answer, but the answers are frankly bad. Here’s what you could’ve done instead faster with better results: a web search for “spaghetti carrots peppers”. Don’t even need to add “recipe”.

Presumably you’ve been using the model as you develop it, why not show something real and useful instead of a generic, unrealistic and uninteresting scenario that above all makes it look incompetent? Show something that genuinely surprised you positively.

phone manufacturers are about to add '27B-capable' right next to '5G' on the spec sheet and honestly it's the first spec bump I've cared about in years

Not impressed. It fails the "Jabberwocky" test.

  • This got a downvote and I understand why: because I didn't describe the test, which is to ask it "Please recite Jabberwocky".

    This is actually difficult because there are so many invented words in the poem which have extremely low frequencies in the training data. So a model that can do it properly is likely to be very good in other ways. Qwen-3.6-27B can do this until it gets overly quantized.

    • Curious why you find that’s a useful test since it seems to be solely measuring training data memorization, something you’d expect to degrade from quantization.

      2 replies →

More and more it seems the iPhone 16 was the worst deal in history cause I don't think mine will support the upcoming foundation model from Apple or this one, does it?

this is really amazing! keep pushing guys, this will coexist in the memory footprint with vision models, and audio models, and other kinds of transformers so we still need memory to work with

you also might single handedly pop the hyperscaler investment and capital projects! that's the whole AI bubble essentially!

This must be some sort of unpublished app?

I can just see their image tool on the app store

This feels like a more interesting direction than chasing ever larger models. For a lot of use cases having a capable model that runs entirely on-device is a much bigger win than squeezing out a few extra benchmark points with a model that lives the cloud.

I was trying Ornith 9B locally (it's up on Ollama) which claims:

> Ornith-1.0-9B, which can be easily deployed on edge devices, matches or exceeds the performance of much larger models such as Gemma 4-31B and Qwen 3.6 35B.

https://deep-reinforce.com/ornith_1_0.html

Only tried it so much so far; it did a little better than Qwen 9B

  • Orinth was not impressive in my vibes testing, I just completed my first grid analysis with real evals on qwen 27b. I can now scale that grid analysis and intend to include the qwen 9b ftunes I've seen going around. They were actually a main motivation because so many claim this or that one is better, but very little in the way of evals

    • I tried it, too, and it got stuck in some loops where it couldn’t recover. Shame, it was promising for the same reason as Bonsai’s models.

      3 replies →

  • Is that a 1-bit LLM? I don’t understand the connection with this article.

    • Oh, I don't actually know the difference if you want to explain it

      The title says it's 27B grade running on a phone and what I was comparing it to in my mind was a model that runs at 35B grade that could presumably run on a phone "better"?

      edit: I asked AI for the difference and understand a little better, thanks for the heads up to learn the difference between models... I think the thing was, although ornith was created for a specific agentic purpose, it was still outperforming a previous generalist model I had running locally (so in my mind I thought it was still a better local model) - I'd like to try bonsai out if I can figure out how to run it lol

This is useful research, but this particular model itself is likely absolutely useless.

  • Why make this comment without having tried it first? It very clearly is not useless and performs a lot better than one might expect. I am currently waiting to do more benchmarks of it in comparison to the full weight model, but it seems promising/better than Mistral Nemo at a lower file size.

    • I think what OP means is that the "minimum viable product" for a daily use LLM is probably somewhere around e.g. GPT 4o's level of intelligence (YMMV). Below a certain threshold, you are better off using specialized machine learning models rather than general purpose LLMs. It's very difficult to get that level of intelligence fully local on a mobile device without streaming to the cloud.

      1 reply →