sethserver.com

All Posts

AI

HIPAA, PHI, and AI: How Not to Accidentally Become a Non‑Compliant Data Processor

Updated: September 17, 2026

HIPAA gets dangerous for builders because it's boring. If your chatbot touches patient-identifying health info, you're handling ePHI-even if you "only pass it through" to an LLM API. This post breaks down Privacy vs Security in dev terms, the vendor/BAA trap, and the controls that matter: BAAs, encryption, least-privilege access, redaction, sane retention, and logging that won't turn into a PHI leak. read on »

AI

GDPR, EU AI Act, and Your AI Stack: Data Minimization for Agents

Updated: September 17, 2026

GDPR and the EU AI Act aren't "legal gotchas." They're design constraints. If you're building LLM agents with tools and heavy logging, data minimization has to live in your architecture: purpose gates, PII stripping before model calls, real TTLs, EU-region routing, scoped tool tokens, and logs you can actually replay and delete. If you can't explain what your agent did with personal data, you can't defend it. read on »

AI

Isolating Tools, Agents, and Tenants So One Prompt Can’t Nuke Everything

Updated: September 17, 2026

Agent failures don't need a hacker. Give one over-privileged agent the wrong tools, add a tiny bug, and you get a feedback loop that refunds money, spams tickets, or leaks data across tenants. This post breaks down the failure modes nobody demos-and the boring isolation patterns that keep the blast radius small: per-tenant tool configs, real execution sandboxes, staging/prod separation, hard policy checks at the gateway, rate limits, and human approval for destructive actions. read on »

Programming

Going From Vibe‑Coding Agents to Production‑Grade AI Security

Updated: September 17, 2026

Built an MCP server in a weekend? Cool. Now check whether you also built an eager little agent with access to your prod database, your `.env`, and a log trail full of secrets. This post is the Monday-after triage: rotate keys, purge git history, put a proxy in front of model + tools, kill "run arbitrary SQL," shrink tool scopes, and treat logs like toxic waste. Shipping the demo is easy. Owning what happens next is the job. read on »

AI

AI in CI/CD: Let the Bot Write, But Make It Earn the Deploy

Updated: September 17, 2026

A three-line YAML tweak from an "AI helper" took down production - not because it was loud, but because it looked boring. This post breaks down how AI sneaks into CI/CD, the failure modes that slip past green checks, and the guardrails that actually help: strict ownership on risky paths, policy gates, staging smoke tests, and a hard "no" on bots merging their own PRs. read on »

Security

Securing Dev‑Facing AI: Code Assistants, MCP Devtools, and CI Bots

Updated: September 17, 2026

Dev-facing AI tools don't need to be "evil" to be dangerous. The real risk is the plumbing: plugins, MCP servers, and CI bots wired to powerful tokens that nobody audits. This post lays out the boring rules that actually prevent leaks-read-only by default, tight scopes, dev/prod separation, human review for writes, and short policies engineers will follow. One pasted config or one sloppy token is all it takes. read on »

Security

Shadow AI: Unapproved Tools, Extensions, and Copy‑Paste Ops

Updated: September 17, 2026

Shadow AI isn't a big vendor deal. It's the tiny shortcuts: a browser extension that "summarizes email," a VS Code plugin that indexes your repo, a chatbot where someone pastes internal docs "just to test." Banning it won't work. People have deadlines. The fix is boring and effective: find what's already in use without blame, ship an approved toolkit that covers real jobs, and set simple guardrails for the copy‑paste zone so speed doesn't turn into a data leak. read on »

AI

Human in the Loop, Not Human as Rubber Stamp

Updated: September 17, 2026

"Human-in-the-loop" isn't a button labeled Approve. If your review UI is a wall of output and your reviewers don't understand the system, you built a rubber stamp-and you'll blame the model when it fails. This post breaks down why review steps collapse, where humans truly must be involved, and how to design review surfaces that work: diffs instead of blobs, plain-language impact, risk signals, and the right kind of friction. read on »

Security

Logging, Monitoring, and Incident Response for AI Systems

Updated: September 17, 2026

If your AI feature ever does something "weird," you won't get a nice stack trace. You'll get a mystery. This post lays out what to log (workflow steps, tool inputs/outputs, prompt + model versions, tokens, latency, correlation IDs), how to redact without building a shadow database of secrets, and which behavior metrics and alerts actually catch trouble. The goal is simple: replay the run, explain what happened, and fix it without guessing. read on »

Security

How to Red‑Team Your Own AI

Updated: September 17, 2026

Demos make agents look calm. Real users don't. This post shows how to red-team an AI agent the way it will actually fail: tool misuse, data leaks, policy bypass, and prompt injection from chats, docs, and tool outputs. You'll get a simple eval harness, ideas for manual attack days, and clear "safe" metrics you can regression-test on every change. read on »

Security

RAG, Vector DBs, and Leaky Knowledge Bases

Updated: September 17, 2026

RAG leaks usually aren't clever. They're a missing tenant filter, a global index, and one "we'll fix it later" endpoint that ships anyway. This post breaks down where cross-tenant retrieval happens, why filtering after search is already too late, and what a secure RAG setup looks like: isolate tenants or enforce pre-filters, tag everything at ingestion, authorize before retrieval, and log exactly what got pulled. RAG is a search system glued to a text generator. If search can see the wrong data, the model will happily repeat it. read on »

AI

King Louie - My Cross Computer, Multi-LLM AI Assistant

Updated: September 17, 2026

LLMs are great until they confidently invent facts and force you to say, "What the crap!?" King Louie is the assistant I built to survive that reality: a desktop, multi-LLM app that shrinks the loop from "ask - copy - run - break - paste error." It adds rule-based model routing, real tools (files, bash, git, web), safe approvals, and even secure mesh pairing so your laptop can dispatch work to your desktop or server. It's chat plus the stuff you do right after chat. read on »

AI

Claude Code's Code Gets Exposed, Whoopsie!

Updated: September 17, 2026

A sourcemap in an npm package exposed the TypeScript source for Claude Code's CLI. Not model weights - just the client. Still, it's enough to see future model names, unreleased features, telemetry (yes, "swearing" counts), and some security checks. The lesson isn't "AI scandal." It's the same old one: if you ship code to the public, assume it will be read. Package like an attacker, and don't publish sourcemaps unless you mean to. read on »

Programming

What Does CI/CD Actually Buy Us

Updated: September 17, 2026

SSH deploys and "git pull" aren't DevOps. They're a ritual that turns one person into a production bottleneck. This post breaks down CI as the boring gate that saves your weekends, CD as the path to repeatable, auditable releases, and why AI-written tests only matter if a pipeline enforces them. If deploys still feel scary, you're missing the seatbelt. read on »

Python

OpenAI Bought Astral - and my fav tool uv

Updated: September 17, 2026

OpenAI is buying Astral, the team behind `uv`, Ruff, and `ty`. I love these tools, but I also get that "someone bought the plumbing" anxiety. This post breaks down why `uv` became my default, what acquisitions tend to break in open source, and the specific red flags (logins, telemetry, AI "help," enterprise splits) that would make me bail fast. read on »

Programming

Terraform and AI are a Match Made in BitHeaven

Updated: September 17, 2026

Terraform isn't hard. It's just painfully specific. Pair it with an LLM and you stop doing "lookup work" (argument names, nested blocks, list vs set) and start doing "review work." The model isn't your architect. It's your fast typist. You still run fmt/validate/plan, read the diff, and keep the blast radius small. read on »

AI

Mistral Moves Closer to My Fantasy with Mistral Small 4

Updated: September 17, 2026

Mistral Small 4 is the kind of model release I actually care about: one thing you can run locally that handles text, images, and code without a pile of routing glue. It's a sparse MoE (119B total, fewer active per token), has a "reasoning effort" speed-vs-depth knob, and ships with a real self-hosting story (vLLM, llama.cpp, Transformers). I'm still going to try to break its multimodal consistency-because models love turning janitors into CEOs-but this is the direction I want: smaller stacks, local control, fewer moving parts. read on »

AI

Neurons Playing DOOM: Why the Python API Matters More Than the Brain Cells

Updated: September 17, 2026

A dish of human neurons learned to "play DOOM" in a week-and the headline is the least interesting part. The real breakthrough is the Python API that makes neuron chips programmable like any other dev platform. DOOM is a tougher benchmark than Pong, but biology still comes with life support, drift, contamination, and brutal costs that don't ship well. The near-term win isn't "wetware replaces silicon." It's hybrid control: silicon does the boring, reliable work, and neuron tissue maybe handles tiny adaptive loops where weirdness helps. My bar stays simple: can it survive outside the lab? read on »

AI

Now Moltbook Gets Bought by Meta

Updated: September 17, 2026

Meta buying Moltbook isn't about "agents posting memes." It's about owning the network where agents discover each other, prove identity, and coordinate safely. Once agents can talk, you get prompt injection, spoofing, leakage, and spam at machine speed. The real work is boring: permissions, constrained actions, readable audit logs, and sandboxes that hold. If you're building here, copy that-not the hype. read on »

AI

The Doctor is in Silicon: AI Isn't Replacing Your Doctor, It's Upgrading Them

Updated: September 17, 2026

Most "AI automation" in healthcare isn't a robot doctor. It's a chance to delete the admin sludge: intake that turns patient-speak into structured notes, chart summaries that cite sources, prior auth packets assembled from the chart, and follow-up messages drafted from approved templates. The real win is boring and measurable: fewer re-typed meds, fewer clicks, fewer denials, and more clinician time for actual judgment. Also, if a vendor can't explain where the data goes and who touches it, assume the "model" is a spreadsheet with humans hiding behind it. read on »

Newsletter

One email, once a week.

Notes on databases, systems, and the occasional strong opinion about Python. No spam, unsubscribe anytime.