On Monday, the company that spent the last year starving llama, poaching researchers with nine figure Zuck bucks and locking its new models behind a Facebook off page just released a free and open- source model under the Apache 2.0 license. It's called Muse Glimmer, a name I believe was inspired by Zuck's fanfiction My Little Pony character, and it's a 30 billion parameter agentic model that's small enough to run on your poverty spec PC. That means you can now have an always on Minis Zuck agent living rent free in your gaming PC, reading your email, managing your calendar, and using your data to overthrow local elections in middle America. And according to Meta, an agent like that needs deep access to personal context, which is a sentence they've been dreaming about writing since 2004, except this time the surveillance can run entirely on your own hardware. In today's video, we'll find out what Muse Glitter Glimmer can actually do. Learn how Meta shrunk it down with distillation, quantization, and speculative decoding, and determine whether Zuck's open- source redemption arc is sincere, or if this is just what losing looks like right now. It is August 12th, 2026, and you're watching the code report. It wasn't long ago when Meta was seen as the champion of the open weight class. Their Llama models slopped up over a billion downloads and created an entire ecosystem of custom models built on top of them. That is until Llama 4 was released and it became the biggest Facebook bust since Justin Timberlake partied with those underage girls. In the release were two models, Scout and Maverick. Initially, they benchmarked well, but that was only because Meta submitted a secret juiced up version to Ella Marina that none of us normies could actually download. And the models we could download had the intelligence of farm animals, which was fitting given their names. And while this was happening, comrade Xi Jinping was dropping open models that were better, cheaper, and released in accordance with Mao Zadong's thought. Then Zuck responded to this humiliation the only way a billionaire knows how. He spent 14 billion on a 49% stake in Scale AI just to acquire its CEO, Alexander Wang. Then he went on a poaching spree for researchers at OpenAI and Google and rebranded the whole operation as Meta Super Intelligence Labs. The first thing his new dream team did was rebuild the stack from scratch and abandon open source entirely. Shipping Muse Spark earlier this year as a closed API only model. That created a new problem though because no sane person would choose Muse Spark over something like Claude or Gemini. So with Wall Street starting to ask questions about the 145 billion of capex being incinerated with nothing to show for it. I assume Zuck flew to his secret compound in Kauaii that the locals say he stole from them and formed a plan which eventually led us here to the launch of Muse Glimmer. It's a dense 30 billion parameter model distilled directly from Muse Spark, their big but closed model using logic distillation, which basically means they had the big model whisper its exact probability distributions into the ear of Glimmer until it learned how to fake it. And in Zuck's manifesto he made after the release, he described distillation as an important principle of how the open- source ecosystem works, which is ironic because every American lab spent the last year accusing the Chinese of doing this exact same thing to their outputs. So how good is it actually? Well, according to the TMBBBS, it clearly beats Gemma 4 and goes bar forbar with Quen 3.6. But my favorite number is the prompt injection benchmark where attacks against Glimmer succeeded 28% of the time. But they're listing that as a win because with Quen it's 40%. And the way they got it to run on consumer GPUs was interesting. At full precision the model would need over 55 gigs of memory. So their first move was quantization where Metampress the weights down to about four bits which shrank it to just under 20 gigs. And the second was speculative decoding which is basically autocomplete for your autocomplete. The way it works is a tiny model called Dlash blurts out an entire block of tokens. Then the big model reads the whole block in one pass and throws out the bad guesses, which they say led to a 3x speed up on a 5090. But to be honest, the actual model is way less interesting than Zuck's manifesto I mentioned earlier. In it, he argues that the real risk in AI isn't rogue super intelligence, but a small handful of companies owning it, and that he doesn't understand why anyone who believes AI will end humanity would rush to build it. which is a reasonable point coming from a guy whose company was just fined $567 million last week for being a public nuisance in New Mexico. He also wants Frontier Labs to hand the US government amid training checkpoints of unreleased models. And he announced a billion dollar fund for the towns willing to host his data centers, probably so they'll stop causing all that ruckus. So, is the redemption arc sincere? Probably not. But an Apache license is an Apache license, and they can't put that genie back in the bottle. Zuck and Wang both say open weights for Muse Spark 1.2 are coming soon, which would mean you could self-host the exact model behind Meta's own coding agent, which is pretty cool. And if you want to test out Muse Glimmer for yourself, you can do it with the sponsor of today's video, Open Router. It gives you a single API to access every LLM, so you only have one endpoint and one build to worry about. You can switch between models manually or use their routers to pick one for you based on a combination of factors like price and accuracy. But my startup still hasn't been able to raise a seed round. So I used open router to rewrite the entire codebase in Rust to attract more investors. Because Rust is a notoriously easy language, I used paro router to route every request to cheaper coding models that could still cargo build. But I eventually ended up splurging for some fancier models for the UI, the dedicated image and speech endpoints to generate a few especially attractive profiles to see the community and get the matches flowing. Open Router is easily the best tool I've seen for discovering new models without getting stuck in 25 different subscriptions. And you can try it out for free at the link below. This has been the code report. Thanks for watching and I will see you in the next one.
Generated algorithmically for Search Engine
Indexing.