Yury's Reading List

Articles and sources curated by Yury

RSS Feed
43 articles
May 08, 2026

Event-driven Programming is Usually a Poor Architecture

Postgres creator and DBOS co-founder Dr. Mike Stonebraker explains why workflow architectures are preferred over event-driven and why AI frameworks are converging on a workflow-style model and supporting durable execution.

May 08, 2026

Implementing a virtual filesystem over Elasticsearch

How to implement a virtual filesystem over persistent data storage, such as an Elasticsearch database, to give AI agents access to shell commands, like grep and cat.

May 08, 2026 Anthropic

Knowledge graph construction with Claude

Build knowledge graphs from unstructured text using Claude for entity extraction, relation mining, deduplication, and multi-hop graph querying.

May 04, 2026

Fine-tuning LFM2.5-1.2B-Instruct with GRPO

Learn what GRPO is and how to fine-tune LFM2.5-1.2B-Instruct with GRPO and Unsloth on OCR receipt extraction for JSON.

May 03, 2026 Ashwin Gopinath

Company Brain, Part 2: Factual Memory

In the first piece, I argued that a real Company Brain needs three kinds of memory: factual memory, interaction memory, and action memory. F...

May 03, 2026 Ashwin Gopinath

Company Brain, Part 3: Interaction Memory

Almost everything important in a company happens in meetings, messages, or emails.That sounds exaggerated until you start listing the actual...

May 03, 2026 LlamaIndex

Why Reading PDFs is Hard

LlamaIndex is a simple, flexible framework for building knowledge assistants using LLMs connected to your enterprise data.

May 01, 2026 Paul Iusztin

What Held Up at 3 AM: One Engineer's RAG Case Study

Ship RAG with Weave CLI: Opik traces, an LLM-judge eval harness, and config-driven swaps across 11 vector databases, chunking, and agents.

Apr 30, 2026 ashwingop

Company Brain: Why Most Companies Have Data But No Memory

One of the hardest parts of any organization is institutional friction. Conversations lose context. Meetings create ambiguous follow-ups. Pe...

Apr 30, 2026 Erik Meijer

Guardians of the Agents

Agentic applications—AI systems empowered to take autonomous actions by calling external tools—are the current rage in software development. They promise efficiency, convenience, and reduced human intervention. Giving autonomous agents access to tools with potentially irreversible side effects,

Apr 30, 2026

A forty-year career.

The Silicon Valley narrative centers on entrepreneurial protagonists who are poised one predestined step away from changing the world. A decade ago they were heroes, and more recently they’ve become villains, but either way they are absolutely the protagonists. Working within the industry, I’ve

Apr 30, 2026

Sizing engineering teams.

I’ve come to believe that most organizational design questions can be answered by recursively applying a framework for sizing teams. Over the past year I’ve refined my approach to team sizing into a bit of a framework, and even changed my mind on several aspects, especially the viability of smal

Apr 22, 2026 Logan Markewich

How LiteParse's Grid Projection Algorithm Parses PDFs

A deep dive into LiteParse's grid projection algorithm — how it extracts text from PDFs while preserving tables, columns, and alignment. Open source.

Apr 22, 2026 jxnl

There Are Only 6 RAG Evals

There are only 6 fundamental ways to evaluate a RAG system. Here's how to do it.

Apr 22, 2026 Eugene Yan

Evaluating the Effectiveness of LLM-Evaluators (aka LLM-as-Judge)

Use cases, techniques, alignment, finetuning, and critiques against LLM-evaluators.

Apr 22, 2026

Demystifying evals for AI agents

Good evaluations help teams ship AI agents more confidently. Without them, it’s easy to get stuck in reactive loops—catching issues only in production, where fixing one failure creates others. Evals make problems and behavioral changes visible before they affect users, and their value compounds

Apr 21, 2026 Doug Turnbull

Metadata: the 3rd kind of retrieval

In search we talk about lexical embedding retrieval But we miss another retrieval philosophy metadata And with LLMs its never...

Apr 21, 2026 Simeon Griggs

Approaches to tenancy in Postgres

There are many ways to slice a Postgres database for multi-tenant applications. Let's look at the three most common approaches and the trade-offs.

Apr 21, 2026

Language Models are Injectiveand Hence Invertible

Giorgos Nikolaou <sup>1,5,</sup> Tommaso Mencattini <sup>1,2,∗</sup> Donato Crisostomi <sup>2</sup>  Andrea Santilli <sup>2</sup>  Yannis Panagakis <sup>4,5</sup>  Emanuele Rodolà <sup>2,3</sup>

Apr 19, 2026 Archika Dogra, Sergei Tsarev, Erich Elsen

Why Your Agents Can’t Read Enterprise Documents — and How to Fix It

The best agents fail at enterprise documents: not because they can't reason, but because they can't read. We’re announcing Document Intelligence: powerful state-of-the-art research turning your unstructured enterprise documents into agent-ready data, at scale.

Apr 19, 2026 Gergely Orosz

The impact of AI on software engineers in 2026: key trends

Our AI tooling survey finds concerns about mounting AI costs, more engineers hitting usage limits, and AI tools having uneven effects upon different types of engineers

Apr 19, 2026 Ryan Lopopolo

Harness engineering: leveraging Codex in an agent-first world

By Ryan Lopopolo, Member of the Technical Staff

Apr 14, 2026 Peter Straßer, Benjamin Trent

Late interaction models: How to scale & optimize in Elasticsearch

Explore techniques to scale late interaction models like ColPali for large-scale vector search in Elasticsearch.

Apr 14, 2026 Peter Straßer, Benjamin Trent

ColPali & Elasticsearch: How to search complex documents

Learn about ColPali and explore how to use it to search through complex documents in Elasticsearch, including tables, figures, multiple columns & more.

Apr 14, 2026 Analytics at Meta

Inside Meta’s Home Grown AI Analytics Agent

Inside Meta’s Home Grown AI Analytics Agent From Hack to Company-Wide Tool The hypothesis was simple: can an AI agent perform routine data analysis tasks autonomously? Data scientists tend to get …

Apr 14, 2026 Matthew Adams

Unsupervised document clustering with Elasticsearch + Jina embeddings

A practical, reproducible approach to unsupervised document clustering with Elasticsearch and Jina embeddings.

Apr 12, 2026

Beyond RAG for Agent Memory: Retrieval by Decoupling and Aggregation

Zhanghao Hu    Qinglin Zhu    Hanqi Yan    Yulan He    Lin Gui

Apr 12, 2026

Benchmarking AI Agent Memory: Is a Filesystem All You Need?

Letta Filesystem scores 74.0% of the LoCoMo benchmark by simply storing conversational histories in a file, beating out specialized memory tool libraries.

Apr 12, 2026

ByteRover: Agent-Native Memory Through LLM-Curated Hierarchical Context

Andy Nguyen <sup>1</sup>, Danh Doan <sup>1</sup>, Hoang Pham <sup>1</sup>, Bao Ha <sup>1</sup>, Dat Pham <sup>1</sup>, Linh Nguyen <sup>1</sup>,

Apr 12, 2026 ashwingop

The Price of Meaning: Why RAG, Knowledge Graphs, and Every Semantic Memory Will Always Fail

TL;DRForgetting and false recall are mathematically, provably, inevitable for any memory system that organises information by meaning. Not j...

Apr 10, 2026 Leonie Monigatti, Danny Williams, Victoria Slocum

An Overview of Late Interaction Retrieval Models: ColBERT, ColPali, and ColQwen

Late interaction allow for semantically rich interactions that enable a precise retrieval process across different modalities of unstructured data, including text and images.

Apr 10, 2026 ashwingop

Geometry of Forgetting: Why Brains and LLMs Fail EXACTLY the Same Way

TL;DRMemory systems for LLMs forget EXACTLY like humans, reproducing the exact numbers from some of the most replicated experiments in clini...

Apr 10, 2026 Clelia Astra Bertelli

Giving AI Agents the Document Understanding Layer They've Been Missing

Most agent runtimes reduce complex documents to garbled text, degrading every downstream task. LlamaParse and LiteParse are installable agent skills that deliver structured, layout-aware, multimodal document parsing — locally or via cloud — so agents can finally reason over real enterprise docum

Apr 10, 2026

Introducing Context Repositories: Git-based Memory for Coding Agents

We're introducing Context Repositories, a rebuild of how memory works in Letta Code based on programmatic context management and git-based versioning.

Apr 10, 2026 George He

Engineering Insights: Failure Modes That Break VLM-Powered OCR in Production

Two hidden LLM failure modes that break agentic document ingestion pipelines in production: repetition loops exhaust resources while recitation filters block legitimate tasks. Learn how to detect, handle, and mitigate both across OpenAI, Anthropic, and Gemini APIs.

Apr 07, 2026 Leonie Monigatti

Database retrieval tools for context engineering

Best practices for writing database retrieval tools for context engineering. Learn how to design and evaluate agent tools for interacting with Elasticsearch data.

Apr 07, 2026

Commits · milla-jovovich/mempalace

The highest-scoring AI memory system ever benchmarked. And it's free. - Commits · milla-jovovich/mempalace

Apr 06, 2026 Dens Sumesh

How we built a virtual filesystem for our Assistant

We replaced expensive sandboxes with ChromaFs, a virtual filesystem over Chroma, to give our docs AI assistant the ability to explore documentation like a developer would.

Apr 06, 2026

Harnessing Claude's Intelligence | 3 Key Patterns for Building Apps

Three patterns for building on the Claude Platform that keep pace with Claude's evolving intelligence while balancing latency and cost.

Apr 06, 2026 Sebastian Raschka

Components of A Coding Agent

How coding agents use tools, memory, and repo context to make LLMs work better in practice

Apr 06, 2026

FUSE — The Linux Kernel documentation

Userspace filesystem:

Apr 06, 2026 trq212

Lessons from Building Claude Code: Seeing like an Agent

One of the hardest parts of building an agent harness is constructing its action space.Claude acts through Tool Calling, but there are a num...

Apr 06, 2026 Terrence O'Brien

A folk musician became a target for AI fakes and a copyright troll

Murphy Campbell has an AI imposter on Spotify and troll claiming copyright of her performances of public domain ballads on YouTube.