Posts

Pinned Post

Who picks the tool

Bild
The previous post ended with an agent that writes a missing tool into a running node. Generate, sanitize, persist, compile, load, execute, about 25 milliseconds after the model answers. I was happy with that for roughly a week. Then I looked at the demo page again and counted the text fields. Tool name. What it should do. Parameters as JSON. I had filled in all three. The gap-finding was real. The deciding was mine. MetaPlannerAgent.run("Publish the article", [ %{tool: "classic_plan", params: %{goal: "Write draft, review, publish"}}, %{tool: "slugify_text", description: "Turns a title into a URL slug", params: %{text: "Hello World"}} ]) Everything interesting in that call sits in the second argument, and I typed it. The agent resolved slugify_text , noticed it wasn't in the catalog and had it written. It never asked whether the task needed a slug at all. So this post is about the agent in front of that, the on...

The tool that wasn't there, so the agent wrote it

Bild
Last time I had two agents doing the same job, one on fixed rules, one asking a language model. The AI agent was free to decide the order of its steps, but it could only pick from tools I had written beforehand. Everything it was able to do, I had guessed in advance. So what happens when it needs something that isn't in the catalog? Normally a polite refusal. I wanted to know what the other answer costs, the one where the agent writes the tool itself, in the running system, without a restart. I built it over a weekend. Getting code out of the model worked after two hours. The rest of the Sunday went into everything that comes after that. A catalog is a list, and lists end Tool calling works the same way everywhere. You give the model a list of functions with their parameters, it picks one, you run it, you hand the result back. That's fine, and it has a ceiling that's easy to overlook. The agent can only ever be as capable as my list. For a support bot with eleven known ...

OpenSpec UI: ein Live-Dashboard für deine Specs

Bild
OpenSpec UI: ein Live-Dashboard für deine Specs Wer mit einem KI-Agenten und OpenSpec arbeitet, kennt das Bild. Im Projekt wächst ein Verzeichnis openspec/ heran: Changes mit proposal.md , design.md , tasks.md und ihren Delta-Specs, daneben die kanonischen Specs der einzelnen Capabilities. Die Absicht hinter dem Code steht damit endlich geschrieben, statt sich in ihm zu verstecken. Nur verteilt sie sich über Dutzende Markdown-Dateien, und der Editor zeigt eben Dateien. Was er nicht zeigt, ist der Zustand. Woran arbeitet der Agent gerade? Welcher Task ist der nächste offene? Und was hat sich verändert, während ich zehn Minuten woanders hingeschaut habe? Genau diese Lücke füllt OpenSpec UI , ein kleines Phoenix-LiveView-Dashboard, das lokal neben deinem Projekt läuft und den OpenSpec-Workspace im Browser zeigt — live, während gearbeitet wird. Der Code liegt offen auf GitLab: https://gitlab.com/public_elixir/openspec_ui Ein Blick statt Datei-Hopping Du startest den Server, gibst d...
Bild
Spec Driven Development mit OpenSpec: ein Praxisbeispiel Die meisten Änderungen an einer Software fangen mit einem Halbsatz an: „Bau mal eben ein Cover in den Artikel ein.” Und dann geht es sofort in den Editor. Unterwegs treffen wir ein Dutzend kleiner Entscheidungen, raten, was wohl gemeint war, und am Ende läuft es irgendwie. Nur weiß hinterher niemand mehr genau, was das System eigentlich können soll. Diese Wahrheit steckt dann im Code, verteilt über viele Dateien, und entfernt sich mit jedem Commit ein Stück weiter von der ursprünglichen Absicht. Spec Driven Development, kurz SDD, dreht die Reihenfolge um. Zuerst wird beschrieben, was das System tun soll. Erst danach geht es an das Wie. In diesem Artikel zeige ich, was dahintersteckt, wie der Prozess mit dem Werkzeug OpenSpec (openspec.dev) konkret aussieht und welche Befehle man dafür braucht. Als roter Faden dient eine echte Änderung aus dem Projekt pdf_finder. Was Spec Driven Development ist Im Kern trennt SDD zwei Dinge, die s...

Semantische Suche für die eigene Bibliothek: eine RAG-Architektur in Elixir

Bild
Angefangen hat das Ganze als schlichter „PDF-Finder". Eine Handvoll E-Books, eine Upload-Maske, eine Volltextsuche, mehr sollte es nicht sein. Aber Sammlungen wachsen, und irgendwann ist ein Archiv so groß, dass man es nicht mehr überblickt. Die Volltextsuche fand nur etwas, wenn ich die exakten Begriffe kannte. Frage ich mich „Wie funktioniert Quantenverschränkung?", das Buch schreibt aber von „verschränkten Zuständen", dann findet die Suche nichts. Was fehlt, ist also keine bessere Suche über Wörter. Was fehlt, ist eine Suche über Bedeutung. Diese Erkenntnis kommt schnell. Spannender ist der Weg von dort zu einem System, das mir eine Frage in natürlicher Sprache beantwortet, mit Quellenangabe, aus meinen eigenen Dokumenten und komplett auf meinem Rechner. Um diesen Weg geht es hier, und um die Design-Entscheidungen, an denen er sich entschieden hat. Die Idee: nicht suchen, sondern antworten Das Verfahren dahinter heißt RAG, kurz für Retrieval-Augmented Generation...

Semantic Search for Your Own Library: a RAG Architecture in Elixir

Bild
It started as a plain "PDF finder." A handful of e-books, an upload form, full-text search, nothing more was planned. But collections grow, and at some point an archive gets so big you can no longer keep track of it. Full-text search only found things when I already knew the exact words. If I ask "How does quantum entanglement work?" but the book says "entangled states," the search finds nothing. So what's missing isn't a better search over words. What's missing is a search over meaning. That realization comes quickly. The interesting part is the road from there to a system that answers a question in natural language, with citations, from my own documents, and entirely on my own machine. This post is about that road, and about the design decisions where it was actually decided. The idea: don't search, answer The technique behind this is RAG, short for Retrieval-Augmented Generation. Put simply, you pull the relevant passages out of yo...

Jido in Practice: Agents in Elixir as Composable Actions

Bild
At Fiatbitcoin, around two dozen small agents run in the background. They fetch prices, read the mempool, score news, compute tax deadlines, send alerts. In the beginning these were all just GenServers and Oban jobs, each built a little differently. That works, but it gets messy fast: every job has its own input and output, its own way of reporting errors, its own idea of what “done” means. Jido is a framework for exactly this kind of work. It hands you a small set of clearly defined building blocks that let you assemble agents from small, testable units. I use it in version 2.1. This article is a tour through real code from the project, and the question behind it stays the same throughout: where does Jido pull its weight, and where does it not. If you don’t write Elixir, a few words up front. A module is a collection of functions. {:ok, value} and {:error, reason} are the usual way to return success or failure, a tagged pair. A GenServer is a lightweight, long-lived process wit...

Beliebte Posts aus diesem Blog

Splitting an ML model and a web app across two BEAM nodes — the technical blueprint

Jido in Practice: Agents in Elixir as Composable Actions

A Foundation Model in the BEAM: On-Chain Anomalies with Google TimesFM in Elixir