AI, automation & information security
Patrick Gawron
Agentic workflows, local-first AI, secure business logic. Daily AI experiments, consulting projects - and a build log of coding sidequests.
Tools
Notes
Quick posts for everything that doesn't need a full article - local LLMs, CSS, keto. One page each.
Browse all notes >Tech news
Google releases EmbeddingGemma 2, a 740M multimodal embedding model
740M parameters (270M text, 170M vision, 300M audio encoders) map text, images, video and audio into one 768-dim space, under Apache 2.0. Small enough for a consumer GPU, and it even runs in a browser tab via WebGPU.
Bartowski reworks GGUF tensor layouts for sharper quantised models
New per-tensor layout maps change how weights are packed in GGUF quants, per bartowski's tests improving quality at the same bit-width - relevant if you pull his quants for llama.cpp.
DeepSeek ships V4.1 Flash, a 552B multimodal MoE model under MIT
552B-parameter multimodal MoE under MIT, context up to 1M tokens. Full weights don't fit one consumer card, but the permissive licence means community GGUF quantisations can follow fast.
YuE2 ships a 3B open-weights music model with symbolic planning
3B parameters, runs on a single consumer GPU, plans song structure symbolically before generating audio - a full text-to-music stack you can self-host instead of renting Suno.