From Obsidian to GraphRAG: What One Month of Building AI Systems Taught Me
Looking back on one month architecting AI systems — from data sanitization and multi-agent coordination to tiered autonomy. The biggest takeaway wasn't technical.

From Obsidian to GraphRAG: What One Month of Building AI Systems Taught Me
TL;DR: Over the course of a single month, I progressed from purging a bloated personal knowledge base to custom-building an orchestration layer for multiple AI Agents, before ultimately re-evaluating when an Agent should be granted autonomous authority. The most valuable lesson wasn't about any individual technology — it was learning to ask the right foundational questions before introducing any new layer of automation into the stack.
Entering September, I had no intention of writing a serialized set of articles. It unfolded organically: every time I resolved an operational bottleneck, the next architectural hurdle surfaced immediately. Sanitizing raw data forced me to reconsider how memory should be structured. Structuring memory led to questioning how multiple specialized Agents should coordinate without stepping on each other. And orchestrating Agents brought me face-to-face with boundary design: when is an Agent permitted to act alone, and when must it pause for human review? This article is an opportunity to step back, examine all four puzzle pieces in context, and extract the overarching architectural principle distilled from a month of real-world building.
A Month in Retrospect: From Raw Data Sanitization to Multi-Agent Coordination
The journey kicked off with a task that, on the surface, had almost nothing to do with modern AI: purging digital clutter. Before letting any autonomous Agent touch my personal knowledge base, I forced myself to perform an aggressive cleanup — permanently deleting nearly six hundred obsolete, duplicate, and low-signal files. At the time, I viewed it merely as routine spring cleaning. Only in hindsight did I realize it was the non-negotiable prerequisite for everything that followed.
The subsequent hurdle was architectural: once the knowledge base is clean, how should it be structured so both human and autonomous Agents can query it effectively? That prompted an in-depth comparison between two fundamentally distinct paradigms in GraphRAG vs. Obsidian: How AI Agent Memory Differs From Your Brain — Obsidian serves human association and cognitive synthesis, whereas a knowledge graph layer powers multi-hop algorithmic reasoning for Agents. Two distinct memory layers, synchronized under strict access controls, never conflating their individual duties.

With structured memory operational, the next challenge quickly emerged: clean, organized data is useless if customer communication channels disregard real-world human behavior. An inbound support Agent must recognize that Vietnamese clients communicate via Zalo rather than email — an architectural realization that forced me to re-engineer everything from response windows to cross-channel identity resolution. And as specialized Agents began running in parallel — one handling research, another generating content, a third managing client inquiries — I found myself performing the very task the system was supposed to solve: holding global context and manually deciding which Agent spoke next. That was the tipping point where I chose to build a dedicated orchestration layer, rather than bolting together off-the-shelf frameworks.
Those four components — clean data, structured memory, communication channels, and orchestration — were never disconnected experiments. They are four tightly coupled layers of a single operational stack, each revealing its structural limits only after the layer beneath it achieved stability.
The Most Critical Lesson: The Best System Is Rarely the Most Complex One
If I had to retain just one takeaway from this entire month, it is this: whenever you introduce a new layer into an architecture, the greatest temptation isn't making it slightly more sophisticated — it is the dangerous assumption that greater complexity equals superior quality. In practice, the exact opposite holds true.
During data cleanup, the easiest way to "automate it quickly" was writing an elaborate automated classification script with endless regex fallbacks. I tried it, and it collapsed — not because the code was buggy, but because within three weeks, I could no longer decipher my own taxonomy rules. When designing Agent memory, the exact same temptation reappeared: knowledge graphs sound vastly more impressive than vanilla Markdown vaults, yet the vast majority of operational queries I needed answered never warranted graph-database overhead. When developing the orchestration layer, I initially considered dynamic peer-to-peer Agent discovery — which rapidly spawned a tangled web of cross-dependencies that only I could manually untangle when an edge case failed at midnight.
Across all three missteps, the pattern was identical: a great system isn't the one that handles the highest theoretical number of edge cases on paper. It is the system that the person responsible for production operations can still clearly comprehend and repair months later, when an unexpected exception crashes at 2:00 AM and no one else is around to debug it.

Four Questions I Force Myself to Answer Before Adding Any Automation Layer
After experiencing that same failure pattern four times in thirty days, I established a strict operational filter before permitting myself to introduce any new automation layer into the stack — not to stall momentum, but to prevent having to tear down next week what I built today:
- Am I solving a concrete, live operational problem, or am I building a theoretical capability "just in case"? Both the purge of 573 files and the custom orchestrator originated from acute, painful friction in daily work, not from reading about a trendy GitHub repository.
- If this layer fails at 2:00 AM, do I have the confidence to open the code and pinpoint the failure within five minutes? If the answer is "I would need a while to remember how this works," that is an immediate signal the component has exceeded my actual operational threshold.
- Is the resulting action easily reversible if it fails? This exact boundary dictates which operations an Agent is permitted to execute autonomously, and which must pause for human confirmation — a principle I dissect in depth toward the end of this piece.
- Am I implementing this because it is genuinely necessary, or because I want the architecture to "look" enterprise-grade? By far the most difficult question to answer honestly — yet the one that consistently protects me from self-inflicted over-engineering.
These four questions carry zero technical novelty. Their value lies entirely in the discipline of answering them before writing a single line of code, rather than post-mortem after an unmaintainable system is already live.

What I Am Testing Next in October
Heading into October, I don't have an over-engineered weekly gantt chart — what I have is an unambiguous testing vector: running all four layers built throughout September concurrently under live production conditions, rather than benchmarking each layer in artificial isolation. Sanitized data, purpose-built memory, channel-aligned communication, and an explicit orchestration boundary — all four look coherent on architectural slides, but only sustained multi-week operational load will reveal where the real systemic friction hides.
I also plan to formalize the autonomy boundaries of each Agent into a reproducible matrix, replacing the heuristic judgments applied during rapid prototyping this month. This is an area I deliberately refuse to draw premature conclusions on — because operational reliability must be demonstrated through runtime logs, not asserted in a summary blog post.

Conclusion
Looking back, this past month was never about learning novel AI frameworks. It was about relearning the exact same fundamental principle four times, across four different tiers of a production stack: simplicity and explicit boundaries will always outperform bloated, full-featured complexity, so long as you remain the engineer accountable for every single component.
The question of operational autonomy boundaries — which decisions an Agent makes solo, and which strictly require a human sign-off — is a topic I explored in detail in Zero Trust for AI Agents: When Does a Human Need to Approve?. If September was about laying the foundation, October will test whether that foundation holds under weight.
Nhận Bộ Thư Viện Prompt & SOP AI Workflow Vận Hành Doanh Nghiệp 2026
Tặng miễn phí Ebook PDF + Notion Template quản lý AI System thực chiến từ Tôi Là Tùng. Gửi trực tiếp vào hòm thư công việc của bạn.
Tặng File Cấu Hình AI Stack Tự Host (n8n + Flowise + pgvector)
Docker-compose sẵn dùng để tự dựng hạ tầng AI của riêng bạn trong vài phút — không cần trả phí SaaS hàng tháng.

Bài Liên Quan

Tại sao tôi từ bỏ viết code truyền thống để chuyển sang Vibe...

Tháng thứ tư với hệ thống AI Agent: khi sự hào hứng biến mất và kỷ luật bắt đầu
