Ivan Fioravanti ᯅ
GenAI/LLM addicted, Apple MLX, Cloud computing, Kubernetes, Technology Advisor, Investor and Co-Founder & Board Member of CoreView.
Unofficial aggregated profile. This creator hasn’t claimed their profile yet.
How to train your own decision model! Unsloth is another level!
UUnsloth AIX· 18h agoYou can now train your own Decision model with our free notebook! 💡 Qwen3.5-4B will generate decisions instead of text on just 8GB VRAM locally. Learn to data prep (state, questions, gold answers), train, serve. Notebook: https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Qwen3_5_(4B)-Decision.ipynb Guide: https://unsloth.ai/docs/basics/train-your-own-decision-model-with-unsloth
TensorFold is pushing like crazy!
AAsh HartX· 8h agoTensorFold 1.0.3 lands with a few cool things. - Living Weights (beta with Nemotron on Metal) and coming to CUDA in 1.0.4 - Nvidia Ampere (beta) for Nemotron, and native Qwen3.8-27B with DFlash2 - Conversation prefix caching - Various bugs fixed and improvements Twelve contributor PRs helped make this happen, so a massive thank you to @niklaslenz_ai and GitHub contributors akol1, jschmied, chaog992, AdrianBinDC and BobClawblaw, and everyone testing, reporting bugs and helping us improve TensorFold. FYI, Living Weights is in very early beta and has sharp edges, but I think I’m on to something. I’d love community feedback to make it better: https://github.com/ashhart/TensorFold/blob/main/docs/living-weights.md @plotarmordev @petruspennanen @MiaAI_lab @volatilemarkts http://tensorfold.dev
Qwen-Image-2.1-Turbo??? Let's try it!
QQwenX· 19h ago🚀 Meet Qwen-Image-2.1-Turbo — create and edit images in just 8 denoising steps! Open weights now available! Built on Qwen-Image-2.1, Turbo is an accelerated checkpoint on the same 7B visual generation architecture. Fewer steps does not mean lower quality: it still generates strong 2K images from text, and supports continued creation through natural-language edits, from adding accessories to changing a scene. Start directly with Diffusers: load QwenImage21Pipeline and the checkpoint’s recommended 8-step sampling schedule is ready to go. 🌐 Prefer an API? Qwen-Image-2.1 Pro and Turbo APIs are now officially live. Choose a hosted API for your application, or use the Turbo weights in your own workflow. 🤖 ModelScope: https://modelscope.cn/models/Qwen/Qwen-Image-2.1-Turbo 🤗 Hugging Face: https://huggingface.co/Qwen/Qwen-Image-2.1-Turbo 🔗 Pro API: https://modelstudio.console.alibabacloud.com/ap-southeast-1/model/market/detail/qwen-image-2.1-pro?serviceSite=international&ref=list 🔗 Turbo API: https://modelstudio.console.alibabacloud.com/ap-southeast-1/model/market/detail/qwen-image-2.1-turbo?serviceSite=international&ref=list What will you create with 8 steps? Share your results with us!
- MMiaX· 17h ago
sparkDash ⚡️ v2.0 is now LIVE!! This is a HUGE release and a COMPLETE revamp! Here's what's new: User Interface & UX - Completely redesigned UI: modern, clean, minimal, yet powerful - Live fleet overview with VRAM, GPU, temps, power and tok/s on every card Run models directly from sparkDash - Add your own start/stop scripts and run your local models straight from the dashboard - Start a model from its Spark's card and see when it's still loading Benchmarks - A dedicated page for each benchmark: Decode, Prefill, Quality and Tool Eval Bench - Built-in quality benchmarks to check how good a model really is: GSM8K, MMLU and instruction following - Tool Eval Bench built in, with ready-made runs - Compare two runs side by side - More accurate Prefill tok/s - Shareable result cards Monitoring - GPU history that survives restarts - VRAM breakdown showing what the model holds, what the system uses, and what's free - Three new pages: Token Totals, Fleet Energy and Activity Total tokens used - Generated, prompt and cached tokens across your whole fleet - Break it down by type, model or Spark - Ranges: 24h, 7d, 30d and all time - Cache hit rate and change vs the previous period - Auto insights: busiest day, busiest model, and which Spark did the most work - Sortable tables and CSV export Fleet Energy - kWh per node, by hour or by day - Average power and peak draw - Efficiency in Wh per 1,000 tokens - Coverage, showing how much of the period was measured - Estimated cost, using the electricity price and currency you set in Settings Activity - Full event history for your fleet: Sparks going online or offline, throttling, and benchmark and model start/stop events - Filter by Spark and event type, with paging - The latest events also appear on the Overview And lots and lots of other cool features! More screenshots below. Get it here: https://github.com/MiaAI-Lab/sparkDash
Clef-Omni is here! It reads the state as text, JSON, images, audio, or video, and returns a probability for every allowed option of every question in a single forward pass. 🚀 Weekend testing!
JJulien ChaumondX· 12h agoand... new clef drop from cloudflare!
Cloudflare/clef-omni · Hugging Face
huggingface.co
We have finally met in person talking about Local AI and data centers! 🙌 Super Marco!
MMarco FranzonX· 1d agoLocal AI team here with the legend @ivanfioravanti We are spreading the verb of the future! If you are at Wave in Turin and you are into local AI say hi!
Amazing thread of Local AI magic!
Kkeys 🧪X· 2d ago‼️‼️‼️ want to see what keys -Mac TensorFold Studio can do? Examples below what it produced👇🧵
GitHub - drowzeys/keys-Mac-TensorFold-Studio: TensorFold Studio for Apple Silicon: Qwen-Image-2.1 (text to image) + MiniMax H3 (video with sound) on MLX with int8 M5 kernels. Text to image to video in one command.
github.com
Just downloaded and activated Underdog. Onboarding is just WOW. Unsubscribed so many unwanted newsletters in few seconds. 💪
UUnderdog AIX· 2d agoToday we're releasing Underdog Saluki 27B Based on Qwen 3.8 27B, Saluki is almost 7x smaller than its full-precision counterpart while it keeps 96% of its benchmark performance and beats Qwen at tool calling It's the best model <8GB that runs in Underdog Saluki is available today under Apache 2.0. Try today in Underdog http://underdog.ai/saluki
Darkbloom is growing! Another company I followed from day one that I bet will become a super star!
KKydoX· 2d agoLocked in the first partnership where Darkbloom will have day-one model support for an upcoming model launch. The first of many!
TensorFold runs on zig now. I should really give this language a try! I bet we'll see a better integration in mlx-serve soon. Great job @ashxhart 🙌
AAsh HartX· 2d agoTensorFold has been rewritten in Zig for Metal/CUDA with some decent performance improvements across the board. Thanks to jschmied, JRaxworthy, BobClawblaw and everyone who sent PRs. Special thanks to @ddalcu for suggesting Zig over Rust and C++ @MiaAI_lab @volatilemarkts @CerebralCoding_ @petruspennanen http://tensorfold.dev
Splash 1.3.0! ⚡️
IInco AIX· 2d agoSplash 1.3.0 is out, and it's about the wait before your local agent starts answering ⚡ If you switch projects and come back, SSD offloading brings the first token down from 19 seconds to 1. New tasks start 1.5× faster, thanks to the Neural Engine. All on a 24 GB M6 Mac.
Hybrid AI is here to stay! 🚀
SSatya NadellaX· 2d agoToday marks a new chapter for Windows, as we bring unmetered intelligence to every desk and every home, and make every PC a place where agents can work securely on your behalf. Some highlights of what we announced: • MAI-Code-1.1 Flash: 137B parameter coding model w/ 256K context window, which is now optimized to run on your PC! • GitHub Copilot now hands off work to local models like MAI-Code-1.1 Flash, helping projects cost a lot less without sacrificing quality. • With Hybrid Intelligence, Copilot can now take action directly on the PC and keep sensitive work on your device. • And with Code in Copilot, you can essentially build any software you need on your desktop, without any cloud token spend, and it’s just super at it. You’re no longer limited to what’s in an app store! Your PC becomes an infinite software factory. • Security is foundational to all this, which is why we are also bringing together Windows and Agent 365 so agents can work within secure boundaries on-device, including MXC a local sandbox for agent execution. Windows becomes your secure agent box! • All this comes to life on a new generation of devices, like Surface Laptop Ultra, powered by NVIDIA RTX Spark. Can’t wait to see what you build with all this.






