Writing
Notes on building production AI.
Multimodal agents, RAG, model selection, and shipping full-stack SaaS — written by Seunghun Lee, the founder behind transcribe.so and goodlisten.co.
Latest
Archive
17 posts
Building a Speech-to-Text Product: The Architecture That Survived Production
The production architecture behind a speech-to-text product: audio chunking, durable job pipelines, multi-model routing, word timestamps, and real cost math.
Fractional CTO vs Technical Co-Founder: Which Do You Need?
A decision guide for non-technical founders weighing a fractional CTO against a technical co-founder: equity vs cash, speed vs commitment, and honest limits.
Shipping an AI MVP in 6 Weeks: The Plan I Actually Use
A week-by-week plan for shipping an AI MVP in six weeks: scope one workflow, build a thin slice, add evals before polish, and get it in front of real users.
How to Vet a Freelance AI Engineer (Questions That Actually Filter)
A practical guide to vetting freelance AI engineers: questions that expose demo-makers, a red-flags table, and why proof-of-work beats portfolios.
How Much Does an AI MVP Actually Cost in 2026?
Real 2026 cost ranges for an AI MVP — solo senior builder vs agency vs in-house — plus the hidden costs (inference, evals, rework) founders miss.
What I Learned Shipping Software at Spotify and Klarna
Concrete engineering habits I picked up at Spotify, Klarna, and a YC-backed startup — and how they shape the AI products I build and operate solo today.
The Next.js + Supabase Stack I Use to Ship SaaS Fast
The exact Next.js, Supabase, and Stripe stack I use to ship and operate real SaaS — pragmatic, no vendor lock-in, and battle-tested in production.
Why Your RAG App Needs Cited Answers
Uncited RAG answers are a liability. Here's why source attribution wins trust, kills hallucinations, and makes debugging tractable — with real examples.
10 Industries Multimodal AI Is Transforming in 2026
A founder-engineer's field guide to multimodal AI in 2026: ten industries, ten concrete RAG and multimodal use cases that ship, plus what actually works.
How to Choose the Right ASR Model for Your Product
A founder-operator's guide to picking an ASR model: weighing WER, language coverage, cost, diarization, and latency across GPT-4o Transcribe, Qwen3-ASR-Flash, and Voxtral.
The Fractional CTO Playbook for Early-Stage AI Startups
A founder-operator's playbook for what a fractional CTO actually does at an early AI startup: architecture, hiring, model and cost decisions, and derisking.
From Prototype to Production: Shipping LLM Features That Don't Break
A founder-engineer's guide to taking LLM features from demo to production: evals, guardrails, fallbacks, cost control, and observability that actually holds up.
Hiring an AI Agency vs Building In-House: A Founder's Framework
A founder-operator's honest framework for choosing between in-house hires, an AI agency, or a solo operator — real costs, fit, and avoiding the prototype trap.
Building goodlisten.co: AI Podcast Discovery and a Creator Studio
How I built goodlisten.co — embedding-based podcast discovery plus an AI studio that turns long episodes into chapters, highlights, and clips. A real case study.
How I Built transcribe.so: Lessons From Shipping a Multi-Model ASR Platform
A founder case study on building transcribe.so: why I route across OpenAI, Qwen, and Mistral, how cited Q&A works, and how I scale long-audio pipelines solo.
RAG vs Fine-Tuning: Which Is Right for Your Product?
A founder's decision guide to RAG vs fine-tuning: when retrieval wins, when fine-tuning wins, when you need both, and the real cost, latency, and maintenance tradeoffs.
Why Your Business Needs Multimodal AI Agents With RAG
A founder-engineer's plain guide to multimodal AI agents with RAG: what they are, why retrieval beats a bare chatbot, and where they create real business value.
Want this kind of work on your product?
I'm Seunghun Lee — I design, build, and ship production AI agents and full-stack SaaS. Tell me what you're building.