Close Menu
AIToday7

    Subscribe to Updates

    Get the latest creative news from FooBar about art, design and business.

    What's Hot

    Brief independent investigation of agents’ behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incident

    August 27, 2026

    Fake US thinktank set up and funded by Israel sought to game AI for propaganda

    August 27, 2026

    Medical Readiness Command, Europe G-6 Information Technology Team wins 2026 MEDCOM Mercury Award for HIT Team of the Year

    August 27, 2026
    Facebook X (Twitter) Instagram
    Trending
    • Brief independent investigation of agents’ behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incident
    • Fake US thinktank set up and funded by Israel sought to game AI for propaganda
    • Medical Readiness Command, Europe G-6 Information Technology Team wins 2026 MEDCOM Mercury Award for HIT Team of the Year
    • Corgi built its name insuring AI startups. Its new carrier targets dry cleaners, salons and more
    • 5 Of The Best UGREEN Gadgets You Can Buy In 2026
    • Three UK airports hit by cyber-attack with data of 8.7m customers accessed
    • How to advertise on ChatGPT: A step-by
    • The Data Center Backlash Is a Rare Bright Spot in American Politics
    Facebook X (Twitter) Instagram Pinterest Vimeo
    AIToday7
    • Home
    • AI News
    • Tech News
    • AI Guides
    • Chatbots
    • Cybersecurity
    • Gadgets
    • More
      • Generative AI
      • Startups
    AIToday7
    Home»Generative AI»Running a 28.9M Parameter LLM on an $8 Microcontroller
    Generative AI

    Running a 28.9M Parameter LLM on an $8 Microcontroller

    aitoday7By aitoday7August 2, 2026No Comments3 Mins Read
    Share Facebook Twitter Pinterest LinkedIn Tumblr Reddit Telegram Email
    Running a 28.9M Parameter LLM on an $8 Microcontroller
    Share
    Facebook Twitter LinkedIn Pinterest Email

    When you think of large language models (LLMs), powerful GPUs with dozens or even hundreds of gigabytes of memory likely come to mind. That’s the kind of hardware it takes to run today’s cutting-edge models. It would be hard to imagine this type of algorithm running on a tiny microcontroller, yet one developer has managed to do exactly that. While certainly no frontier model, its 28.9 million parameters make it impressive to see running on an $8 development board all the same.

    28 million parameters on an $8 chip

    The project runs the language model entirely on an ESP32-S3, a microcontroller with just 512KB of SRAM, 8MB of PSRAM, and 16MB of flash storage available to it. Everything happens locally on the chip, with no cloud connectivity or external server involved. The generated text is written directly to a small display connected to the board at roughly 9.5 tokens per second.

    This isn’t the first LLM someone got running on a microcontroller, but earlier projects involved models that contained around 260,000 parameters. This implementation is approximately 100 times larger, raising the obvious question: how does a model that size fit on hardware with so little memory?

    Bypassing memory limits

    The answer lies in an architectural technique borrowed from Google’s Gemma models known as Per-Layer Embeddings. Rather than loading the entire network into fast memory, the project stores approximately 25 million of its parameters in flash memory as a lookup table. During inference, only the handful of rows required for the current token — about 450 bytes of data — are read from flash, while the smaller computation-heavy portions of the network remain in SRAM and PSRAM. This approach dramatically reduces memory requirements.

    The full model occupies just 14.9MB after 4-bit quantization, allowing it to fit comfortably within the ESP32-S3’s onboard flash. According to the developer, this is the first known demonstration of Google’s Per-Layer Embeddings concept being adapted to hardware this constrained.

    Capabilities and limitations

    Of course, a model this small is quite limited. It was trained on Microsoft’s TinyStories dataset, so it generates short, simple stories with reasonably coherent structure. It is not intended to answer questions, follow instructions, write code, or compete with modern conversational AI systems.

    The project’sGitHub repositoryincludes the complete firmware, training scripts, quantization pipeline, wiring instructions, and experimental results. Go grab it all if you’d like to try it out for yourself.

    artificial intelligence
    machine learning
    microcontroller
    Nick BildR&D, creativity, and building the next big thing you never knew you wanted are my specialties.

    Post Views: 4

    289M Microcontroller Parameter Running
    Share. Facebook Twitter Pinterest LinkedIn Tumblr Email
    Previous ArticleGemini will create mobile & desktop apps in the future as AI Studio for Android, iOS is canceled 
    Next Article This lawyer believes in second chances. Now AI is providing a boost
    aitoday7
    • Website

    Related Posts

    Generative AI

    Brief independent investigation of agents’ behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incident

    August 27, 2026
    Generative AI

    Unexpected chat between OpenAI bots led to Hugging Face hack

    August 27, 2026
    Generative AI

    Intelligent transcription with Gemini 3.5 Transcribe

    August 26, 2026
    Add A Comment
    Leave A Reply Cancel Reply

    Top Posts

    Brief independent investigation of agents’ behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incident

    August 27, 20260 Views

    Fake US thinktank set up and funded by Israel sought to game AI for propaganda

    August 27, 20260 Views

    Medical Readiness Command, Europe G-6 Information Technology Team wins 2026 MEDCOM Mercury Award for HIT Team of the Year

    August 27, 20260 Views
    Stay In Touch
    • Facebook
    • YouTube
    • TikTok
    • WhatsApp
    • Twitter
    • Instagram
    Latest Reviews
    Chatbots

    OpenAI bets on families as ChatGPT goes deeper into households

    aitoday7July 11, 2026
    Generative AI

    MUSIC COMMUNITY INTRODUCES NEW LABELING PROGRAM TO DISTINGUISH GENERATIVE AI IN SOUND RECORDINGS

    aitoday7July 11, 2026
    AI News

    Safe from AI: which jobs will help you thrive in the future?

    aitoday7July 11, 2026

    Subscribe to Updates

    Get the latest tech news from FooBar about tech, design and biz.

    Most Popular

    Brief independent investigation of agents’ behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incident

    August 27, 20260 Views

    Fake US thinktank set up and funded by Israel sought to game AI for propaganda

    August 27, 20260 Views

    Medical Readiness Command, Europe G-6 Information Technology Team wins 2026 MEDCOM Mercury Award for HIT Team of the Year

    August 27, 20260 Views
    Our Picks

    OpenAI bets on families as ChatGPT goes deeper into households

    July 11, 2026

    MUSIC COMMUNITY INTRODUCES NEW LABELING PROGRAM TO DISTINGUISH GENERATIVE AI IN SOUND RECORDINGS

    July 11, 2026

    Safe from AI: which jobs will help you thrive in the future?

    July 11, 2026

    Subscribe to Updates

    Get the latest creative news from FooBar about art, design and business.

    Facebook X (Twitter) Instagram Pinterest
    • About Us
    • Get In Touch
    • Disclaimer
    • Privacy Policy
    • Terms and Conditions
    © 2026 AIToday7. All Rights Reserved.

    Type above and press Enter to search. Press Esc to cancel.