From the Introduction: The Advice Gap
A free preview from "Databases in the Era of AI"
Most database advice clusters at two extremes, and neither one is aimed at you.
On one end you have vendor content — written to make you think you need a specific product. On the other you have war stories from teams running thousands of nodes, sharing hard-won lessons about problems you will not have at your scale. The middle, where a five-person team is trying to make one good call, gets almost nothing.
You've read the blog posts. You need a graph database for recommendations. A time-series database for metrics. A vector database for search, obviously, because it's an AI feature now. Follow that advice to its logical end and you're running five database systems with three people, two of whom also write the application code. Each new system is another thing to back up, monitor, and explain to the next hire.
I once wrote custom BI software in C++ for AT&T, complete with my own encryption algorithm to pass database credentials between machines without an external library. It was clever. I was proud of it. An off-the-shelf enterprise BI tool would have done the job better and saved everyone months.
I still get that itch when I see a hard problem and want to invent something beautiful for it. What's changed is that I catch it sooner — usually before I've written the encryption algorithm.
The part people don't want to hear when they're excited about shipping their first AI feature: your data layer mostly stays the same. The Postgres database running your billing, your users, your core product keeps doing what it does. AI adds new types of data on top of that. It doesn't burn the old stuff down.
What's new is a short list. Embeddings. Prompt-completion pairs — the text you sent to a model and what came back. Retrieval logs. Model metadata: which version, what temperature, how many tokens, how long it took. Look at what those things actually are: text, numbers, timestamps, arrays of floats. None of it is exotic. Your relational database already handles text and numbers. It can be taught the float arrays with one extension.
An embedding is a fixed-length array of floating-point numbers. The reason people think this needs its own database is that you can't search it with a normal WHERE clause = something. But "different kind of index" and "different database" are not the same problem. Postgres handles this with an extension called pgvector. Your embeddings sit right next to your users and invoices. One backup plan, one connection pool, one thing to operate.
This is just a preview. Get the full Introduction plus Chapters 1 and 2 free when you sign up.
Get the first 3 chapters (Introduction + Chapters 1-2) instantly delivered to your inbox.
Get Free ChaptersAll 13 chapters + architecture recipes + migration playbooks + the reference card. Start reading tonight.
Buy Now - $4.77