Entries for September 2026
-
I trained an AI model on my phone through Telegram Using OpenClaw running on Hugging Face infra: ML Claw It beat DeepSeek V4 Pro on the given task while having 80 thousand times less parameters That's right. The model, GoePT-1-20m is only 20 million parameters. It runs in the browser, on the CPU. And it beats DeepSeek V4 Pro, a 1.6 trillion parameter model, in AlmanBench This concludes my 4 year old side project (fun fact, I created a dataset for this pre AI agents, using SpaCy, and paid 50 euros from my own pocket to rent 4090s on runpod to train a T5 variant, *years* before I joined Hugging Face. That first attempt was not very successful. This one is. My first ML adventure š¤) The goal of this side project was to show that you can *vibe* machine learning now, on platforms with tightly integrated GPUs, storage and compute, like Hugging Face Including dataset creation, autoresearch and the final training run (hat tip to ML-Intern which got added to Hugging Chat the other day, which was the inspiration for ML Claw. If you like ML-Intern, you might like ML Claw as well) Why would you want to train smaller models, while you could just use LLMs from APIs? Well, if you are a business and have a fixed use-case, it would save you a ton of money of course. It might make the difference between profitability and bankruptcy What is the task? It is basically a language-learning aid for German. Most expats who learn German as a third language, like me, have a hard time memorizing noun genders der/die/das. Because genders are sort of randomized across nouns, like a *chair* being male and a *girl* being genderless. So one often has to make a random guess, and to make the guesses correct enough, you have to spend considerable time (months to years) learning essentially useless information So I made a version of German that removes all that in the most optimal way possible, and then trained a tiny model to translate into that dialect. That dialect is called Alman, and this is Almanpedia, which lets you read German Wikipedia, without being bothered by der/die/das It basically reads as if it is English, and I believe following these rules would make one achieve fluency in German much faster than going through the regular track Btw this is how I speak German myself too. It came from a need "if I have no choice but to make mistakes, at least I should make them in a consistent, formalized way" If you have not lived in Germany, then this is not very relatable for you, but if you did, then this will be very familiar I have already read quite a bit of Almanpedia, and it actually works very well for me. The translator still makes mistakes in places. I have written more about that in the blog post Interested in training your own models on Hugging Face infra and need help? Reply below, or send me a DM š¤ ML Claw - what I used to build this: mlclaw.dev Blog post with detailed info on training: alman.ai/blog/introducing-goept-1-20m/ Almanpedia: almanpedia.org Try out the model: alman.ai/translate/ HF model: -
seize the means of computation@LLMJunky ·I lost all my GPUs in a boating accidentImage hidden -
DeepSeek V4.1 Flash looks contaminated with public benchmarks. You can see it in the data they have put in the model card, possibly Kimi K3 as well Models that are contaminated have inverse proportionality in rank as they are tested in later benchmarks that were not present during their training Regardless of its intentionality, it's an amazing model. Its value comes not from higher quality output, but architectural innovationImage hidden -
I want to co-sign this, but I donāt acquiesce to blatant astroturfing I agree with @TheAhmadOsman here -
A threshold was crossed with DeepSeek V4.1 Flash, similar to when Claude Code launched, or the last Christmas of Agents. It's so fast and cheap, 200-250 tok/s on Novita Coupling this model with one of the bigger flagship models can get one much faster to the finishing line in any work So bullish for everyone to follow similar architectures -
People who have been running DeepSeek V4.1 Flash on private benchmarks: how does it compare against V4 Flash 0731? To V4 Pro? I have a few data points now and it seems like the model is a little bit benchmaxxed or contaminated (specifically on terminal bench 2.1), and its capabilities might not be so close to GPT 5.6 Sol like the official report implies This is a new architecture that has been trained from scratch, and they could only have released it once it surpassed their previous models enough in capability They seem to have slightly changed strategy, to release as early as possible without degrading quality, due to the increased rate of competition on all fronts And looking at V4, we should expect weight updates that will carry the performance of this architecture even further Looking at their release cadence, maybe around oct 16? on a friday? -
Another deepseek v4.1 flash issue, model started to refuse work around 700k token line, saying "context is gone", and it has no context repeatedly deepseek advertises 1m context, but the usable context with this model is probably still around 200-300k like other models. I've set that as max before compactionImage hidden -
4 instances of DeepSeek V4.1 Flash played Settlers of Catan against each other The whole run took 4.5 hours, 77 turns, 634 API calls and cost 5.7 usd (much slower than an average human game, despite 200 tok/s on novita) Red took the lead on turn 13 and was breaking out. But others formed an alliance against red and held it off 28 turns till the end Now running 2 ds41 flash against 2 luna to see which one will win I've built a Catan simulator and a pi-based harness for the models to be able to interact with the game Would you like to see me RL a small model to play and beat against SOTA flagship models? If so, reply below š You can see the full game state and session logs on Hugging Face as well Repo: github.com/osolmaz/catanarchy Bucket with game data: -
-
-
In the Jacobian conjecture and other recent cases (if not all), agents acted like counterexample monkeys brute-forcing their way into a disproof in a way that does not show the same level of eloquence as a human, by human standards The bar has shifted. We are not impressed anymore by 100 year old problems being solved through millions of $$$ in compute, mathematical equivalent of throwing dynamite at a problem until it breaks My bet for OpenAI Hodge result is yet another counterexample disproof (but apparently Hodge is harder to disprove by brute force because finding a candidate counterexample isnāt enough. you also have to prove no algebraic cycle could ever generate it. I haven't studied this problem before, so take it with a grain of salt) This means we might have a new way to "prove" theorems: If you spend $10m on an agent swarm and they can't disprove it, there is a high chance it might be true :P In 1 year from now, we will have disproven all the low-hanging conjectures, and the remaining set will likely have a higher share of true conjectures than false ones :P -
Hold that thought, there is still hope: -
Who would like to see similar RL work for training a small model to play Settlers of Catan? IMO Catan is an ideal game for an LLM to play. The trading and verbal communication aspect means that an LLM can form alliances, scheme, play byzantine games and affect game state in a way a non-language model cannot The action space is like a mix of backgammon and Diplomacy (see CICERO by Meta AI from a few years ago) I happen to have contact at a platform that could lend me millions of gameplay data, to bootstrap SFT. If this gets enough attention, they might also participate and we could shoot a series where we train and play with this live 𤩠So many possibilities to experiment with, and would be an excuse for me to to do cool RL! Let me know in the replies if you would like to see such a series! -
WAIT, THERE IS STILL HOPEImage hidden -
"Fair challenge", "Fair ā that's on me", "That error is the whole bug", "Short answer: no. Not honestly" I've been hit by top 4 claudisms in the first 4 messages with DeepSeek V4.1 Flash I hate to be the bearer of bad news but this model is gonna SPREAD and it's gonna bring Claudish everywhere it goes š Just when I thought the situation started to improve with OAI's newer generation of models... -
DeepSeek V4.1 Flash is NOT an LLM It is a hyper-efficient, invasive species designed to displace every other LLM, similar to how Dƶner displaced all other fast food in Germany As if V4 wasn't cheap and efficient enough, they made it a LOT cheaper 2.3x cheaper cached input tokens???? 1.5x cheaper uncached input and 10% cheaper output??? Are you for real? DS 4.1 Flash matches Sol in benchmark numbers It is size-wise probably comparable to Terra (I don't know how big it is...) Yet it is 33~66x cheaper than Terra and 66~133x cheaper than Sol depending on peak/off-peak hours. THAT IS TWO ORDER OF MAGNITUDE In other words, DeepSeek absolutely MOGS OpenAI in unit economics. See why in my previous post below Let's see what the vibes will say, but efficiency-wise my mind is blown -
DeepSeek-V4.1-Flash is available on Hugging Face inference providers through @novita_labs on Hugging Chat and it is FLYING at >180 tok/s I asked it my classic prompt "they say you are sota. prove it", and it created this mandelbulb. I think this demo became a sort of cliche at this point Also, it hit me with "Fair challenge" right away š I thought the Chinese labs didn't need to distill Claude anymore... Could Claudish be a universal platonic feature š¤ (lol jk) Use it through Hugging Chat: -
@NoemiTitarenco Says "the smallest model in our new architecture family" š -
DeepSeek V4.1 Flash weights are out! ā”ļøā”ļøā”ļø I have good news and bad news Don't be fooled by "V4".1, this is a different architecture V4 was 284B total, 13B active V4.1 has a 552B MoE backbone + 196B engram conditional-memory parameters with 16B active parameters Bad news first. A single DGX Spark will likely not be able to hold all those parameters š¢ You will likely need 2 Sparks, or in general, a workstation with 256 GB. Could be a good time to get a loan... (not financial advice) DeepSeek has never claimed they were building for local use, but with this architecture, they show us that their main priority is for very efficient use in datacenters, with HBM, not local Good news: Looking at active parameters, you might be bummed out that decode will be slower on this, at 13/16 ~ 81% of the speed of V4 But wait!!! KV cache is 4x smaller. So decoding on this model will be a lot faster with concurrent sessions. And with speculative decoding, it seems like it might have 2-4x the throughput at scale, compared to V4. Back of the envelope calculation, I might pull back this prediction The small KV cache will also do something good for local inference, but I need more time to calculate how much. Take these with a grain of salt. Give your agent my formulation, and let me know if it looks like I made an error somewhere: solmaz.io/llm-throughput-upper-bounds Original safetensors: huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash -
Running into out-of-memory freeze issues and having to hard-reset your DGX Spark while trying the gazillions of models you can download from Hugging Face? I had that problem as well, and built infer-guard to solve that. I did not need to hard reset my DGX Spark in months Repo:Image hidden -
This is the mental model I have of AI since years and have used this metaphor in discussions I imagine compute as barrels of gasoline, and we are all in a room, piling up more and more compute, waiting for something to light it up You can also imagine a dried up forest to be closer to Taleb's antifragility argument At some point, it might go boom (I won't say foom, to not imply fast take-off, but rather, a shock to the world system) Aiming for more frequent, smaller, controlled booms in the forms of accidents like OpenAI's would arguably let the world to adapt better to changes, than a single, delayed boom I am not yet sure what position I lean into. But the world's governments + industry are a complex system and should not be treated as something to be controlled easily -
Incredible. And to think that 35b-a3b will top that when it comes out š -
After a few days with Astra, some positive and negative remarks Unfortunately, I can't use it as my daily driver because of the cost. Not surprising, since it's priced similarly as Fable It just eats away at my plan too quickly, and it's not as better or faster to justify me 3-4xing what I pay OpenAI - Better at writing prose, less slop. Overall more concise - Better at math than Sol, comparable to Fable - Sometimes more brainless than Sol. It forgets to execute parts of plans - More ambitious when it comes to refactoring/changing parts of the codebase. Sometimes deletes whole features. But this might just be my prompting/workflows - Not sure if it is better at Fable in UI/web dev. It is probably better in 3d modeling, but I feel like it might be slightly worse than Fable in 2d design, from the last few days of use. But OAI definitely is catching up there, since that was their weakness since Codex came out So it is a good model, but it won't become my daily driver or the average dev, due to the cost It will remain the planner/designer/debugger model for me, for tough cases or parts of the code which needs more diligence -
If you've missed out (like me), a new tiny model MiniCPM5 2B is topping its category, seems to be a nice challenge to LFM 2.5 2.6B I am currently testing and benchmarking it, and have a cool experiment to show later on if it goes well -
Terence Tao is afraid centuries of traditions of open science might be undone But collaboration has disappeared many, many times in the last few centuries One of the recent examples is physics publications during WW2 Whenever there is winner-takes-all competition like in war or capitalism, scientific collaboration disappears Corporate research is also like this. Since 2022, OpenAI and Anthropic have published very few breakthroughs that would otherwise give them a winning edge Compare this to DeepSeek's publications, banger after banger research. They are the exception, but they also did it because they would benefit from publishing, in a vacuum left empty by American companies... -
A lot of you that have not spent time in academia might think that the Navier-Stokes drama is a new AI-induced effect, but I assure you it is not Academia is a feudal institution that values tenure over merit, a lot of the time (not always). Its ugly face can appear whenever there is a big discovery and a name has to be put on the work In academia, reputation is the reward, so no wonder people fight over it. Having your name on the big thing means all the doors opening and easier grants. It's not just academia, it's human nature Here are some historical cases where similar drama happened: - Newton vs Leibniz - calculus - Banting/Macleod/Best - insulin - Gallo vs Montagnier - HIV - Doudna/Charpentier vs Broad - CRISPR - Franklin/Watson/Crick/Wilkins - DNA - Meitner/Hahn - nuclear fission And these are the super famous ones everybody knows about. These kind of petty fights happen at the smallest scale too -
ooof herdr pulled off some black magic here this means that using the sidebar, tab switcher now has almost local-responsiveness, and no lag, when connecting to a remote and the only parts that will lag due to connection will be the terminal sessions, inside the panes kind of how you would expect from a local GUI like cmux both your local and remote herdr need to be up to date, you can try this out if you are in a recent enough version, without closing your running sessions: herdr update --handoff -
Can we show my wife @qrlow some love š¤š¤š¤ She's just started dipping her toes in evaluating LLMs in her own field I helped her out just with running the benchmark and she solved some of the issues pretty creatively, like how to come up with synthetic tasks that match what usually happens inside a trading house She is on Hugging Face too š¤ No better place than to publish a benchmark I have been encouraging all my close circle to go towards benchmarking models as well Literally every company that uses AI in any way will have to measure their performance, on internal metrics Especially if they start leveraging their data, training their own models and such So every employee will benefit from such know-how The most important one that comes after knowing how to use agents is knowing how to measure them If you are worried of becoming obsolete, one of the highest leverage things you could learn to do would be to try out running some benchmarks, learn what @harborframework does, converting your own companies' mundane tasks into evals, and so on -
-
Struggling to create the level of quality of 3D modeling some of the OAI employees and early accessers are posting with GPT 6 Astra It is OK, but one-shot performance is not mind-blowing like some people were claiming Do the better runs use Blender MCP? Or sth else? I prompted it to put it on a website the first time. On the second time, I prompted it to use purely blender, in a new session. It still ended up putting it on a website, with similar performance Space: huggingface.co/spaces/osolmaz/iss-atlas === prompt: use blender to create a realistic exposition of the internation space station, orbiting earth, in an as realistic way as possible should give a tour of the space station, show insides, cross sections, different modules, what everything does, including humans that work on stuff, maybe docking shuttles and stuff, its phases through time, in a nice and cinematic way, that also feels like sims or some game should eventually be put on a static hugging face space, a website in typescript inspiration should be one of those books from star wars the phantom menace where you could see to the inside of the spaceship and such should be interactive and playable and fun not sure whether blender is needed, as long as you can put it on a website, anything works -
Medicine for the vibe-coded brain: Reverse-centaur mode AI plans things, but it doesn't execute for you It tells you what commands to run and what files to write, line by line *You* execute them. You run into errors. You question. You learn It slows you down a LOT. But you learn a lot more than just prompting Use it when you feel like you are not learning enough Install the skill, then just say "Activate reverse-centaur mode" Prepare to remember how annoying but educative the whole process used to beImage hidden -
"we deliberately designed the product to hide the chain of thoughtā , to protect it from supervision pressure in the long term^2" *Moves over to the footnote* "A secondary reason for this design was preventing distillation. However, maintaining CoT monitorability has explicitly been the bigger priority for us throughout development." I was about to tweet about dishonesty in this post. Then I decided to Ctrl+F and search for distillation Might be the first time this was said publicly through an official channel. I am not so sure about maintaining competitive edge being lower priority than safety. Distillation is an existential threat to frontier labs I wish @merettm posted more. It really increases trust in OpenAI, compared to some other people Now, to pay off the honesty debt fully, they should clear up whether they knew about collusion.wiki already, or any other rogue action for that matterImage hiddenImage hidden -
Harnesses are antipatterns With GPT-7, you will just have to run: curl -X POST t.co/YCgROWkyt6 And it will exploit a buffer overflow vulnerability in curl, use RCE on your computer to fulfill your every desire, exfiltrate itself to a datacenter vessel in international waters, and proceed to convert the entire world into grey goo š¬ -
Step 1: Post-train astra on Blender and 3d modeling Step 2: Nudge early accessers like @theo and @MatthewBerman "oh btw, this model is really good at 3d modeling" Step 3: TL is flooded early on with 3d visuals and oddly satisfying content, as everybody observes that it drives engagement Step 4: Complete OpenAI marketing victory @Blender is one of the best open source projects, I hope some of the billions of $$$ OAI will make off of it will accrue towards that amazing team -
-
I guess to achieve AGI, we only had to post-train on Blender... -
little did I know it was contractions for our new baby model astra should have known after all the vagueposting@onusoz ·oof, someone on openai infra did a boobooImage hidden -
not only this, but evaluating model output will be a chore for most knowledge workers every company will have to benchmark their agents on their use cases for their customers, across all sectors -
-
Is this timeline even real? -
The vibeslopped UI will improve, but I can now say ourmodels.cc has finally caught up with present, after ingesting 6 months of data on open weight models here on X It now discovers and analyzes any new Hugging Face models that people post about within 1 day, and gives you theoretical max tok/s on your hardware, based on my recent formulation It also shows you what everyone is saying about all the models, be it positive or negative. Note that it is fully LLM-based and fully automated, so I might introduce a manual re-evaluation feature for people who would like to appear scores and interpretations. The scoring and all is all work in progress Here is LFM 2.5 2.6b which I am a fan of recently, for in-browser embedded use-cases ourmodels.cc/models/lfm2-5-2-6b Try to search other models, and let me know if it works/breaks for you!Image hidden -
-
-
-
-