Tôi Là Tùng
Back to Blog

Dual-Engine AI Agents: Lessons from System 1 and System 2

Dual-Engine AI Agent lessons from real project code: System 1 routing, System 2 reasoning, shadow tests, honest cost measurement, and key and token protection.

Dual-Engine AI Agents: Lessons from System 1 and System 2 | Tôi là Tùng, toilatung, Nguyễn Thanh Tùng, Tùng Sóc Sơn

TL;DR: Dual-Engine separates narrow classification in System 1 from planning and content generation in System 2. Reviewing the code in my agent ecosystem highlights the importance of controlled handoffs: distinguish shadow evaluation from execution authority, identify fallback traffic, measure the whole workflow, and keep authentication separate from model judgments. Integration code alone does not establish production performance or security.

“Write a post” and “send the draft again” may contain overlapping content keywords. If a router confuses them, the system creates another article instead of finding an existing document. A more capable model may help, but the design still needs to define what the model decides, what evidence allows the next step, and who can authorize an action.

This is how I approach Dual-Engine architecture through the system I am building. This September 23, 2026 update checks the CRM and Telegram gateway source code in the agent ecosystem. A call site in source code does not prove a deployment is live or establish performance on production traffic. That distinction separates the implementation evidence below from proposed improvements and measurements still needed.

What are System 1 and System 2 in a Dual-Engine design?

Dual-Engine divides work between narrow decisions and tasks requiring broader context. System 1 selects a label, identifies intent, or evaluates a focused condition. System 2 plans, synthesizes documents, and generates longer outputs. These names are a design analogy, not a claim that a model contains two physical brain regions.

The TypeSafe documentation describes Jev as accepting state and typed questions. Choice selects an option, Noul evaluates a proposition, and Score evaluates a rubric. These primitives suit well-scoped decisions. A correctly typed answer can still be semantically wrong, so the application must validate values and decide how to use them.

For me, this follows Director Mindset: define the responsibility of each step before choosing its tool. Deterministic checks such as access permissions, payload limits, and order status comparisons belong in application code. Calling a model adds little value when the answer already follows an exact rule.

Dual-Engine AI Agent architecture separating System 1 and System 2 responsibilities | Toi La Tung, toilatung, Nguyen Thanh Tung

What has actually been integrated?

The source review shows integration beyond an isolated reference module. However, different call sites use the resulting decisions in different ways.

Reviewed componentWhat the source establishesWhat it does not establish
CRM routerContent, department, and agent classification with a heuristic fallbackThat every request actually uses Jev
CRM dispatcherA shadow classification attached for comparisonReplacement of the existing department decision
Inbound webhooksCalls for input screening and routingVerified deployment, authentication coverage, or execution safety
Telegram gatewayRouting calls for text and media tasksA verified production performance benchmark

This distinction matters more than an “integrated” label. Shadow mode permits observation without giving a new engine authority to replace the existing business decision. A path that already consumes the result to choose an agent needs additional scrutiny because a misclassification can affect subsequent steps.

One detail is particularly useful: the dispatcher's shadow call is awaited before the response returns. It does not change the department decision, but it can still add waiting time. Shadow evaluation therefore does not automatically mean zero performance impact. That lesson follows from the control flow; it does not require an invented savings percentage.

Lessons from routing, timeouts, and fallback

A timeout is a configured waiting limit, not measured latency. The CRM module configures 2,000 ms for routing and 1,500 ms for input screening. On a sequential path, those waits can accumulate before content generation and result handling begin. Model-level timing cannot be substituted for end-to-end workflow latency.

Fallback preserves a processing path, but may change decision quality. When configuration is missing or a call fails, the router uses heuristics. An engine field distinguishes the paths. Counting successful responses alone could make the new integration look healthy while a large share of requests actually use fallback.

A score assigned in code is not measured accuracy. Some heuristic branches return fixed confidence values, and answer handling includes defaults for missing fields. Operational reports should distinguish model scores, application defaults, and correctness against human-reviewed labels. A convenient number in a response does not demonstrate calibration for the application.

Routing precedence deserves explicit test cases. An evaluation set should contain creation requests, requests to retrieve a link, status questions, and mixed intents. “Send the video draft again” should test retrieval separately from the presence of the word “video.” This is a proposed evaluation case, not a claim of a successful customer experiment.

How should the handoff boundary work?

A useful boundary checks authority before model evaluation and validates the action before execution. The following is a recommended design, not a complete representation of every current endpoint.

Receive request
  → Authenticate the sender and establish the data scope
  → Apply deterministic rules and minimize the supplied context
  → Use System 1 for a narrow classification
  → Validate the answer, confidence, and risk policy
  → Use System 2 for synthesis or planning where needed
  → Review the output and obtain approval for important actions
  → Execute within the permissions already granted
SituationDesign decision
An exact rule determines the answerHandle it in code without another model call
A narrow, low-risk classificationEvaluate System 1 against the current baseline
Ambiguous intent or missing informationAsk for clarification or escalate for analysis
Missing security-screening resultBlock or queue for review according to policy
Publishing, payment, or permission changesEnforce authorization and a separate approval gate

A heuristic can be a reasonable fallback for content classification. It should not automatically turn a failed security check into permission to proceed. The reviewed modules still contain branches that treat an input as safe when no regex matches, and low-risk defaults when an answer is missing. These are limitations to improve, not evidence that a complete fail-closed policy is already implemented.

Security: keep keys and tokens out of both articles

A prompt-injection classifier is a supporting layer. It does not prove that the sender is authenticated, determine document permissions, or replace isolation between accounts. A model's “safe” label never grants tool access.

Public examples should use synthetic inputs rather than copied configuration screenshots or raw logs. API keys, access tokens, session cookies, signed access values, and links carrying authentication parameters must be removed. Configuration variable names can help technical documentation, but credential values have no place in this article.

Review both language versions, metadata, code blocks, linked URLs, rendered HTML, and structured data. Translation does not sanitize secrets: a sensitive string may survive unchanged in the English version. Screenshots and attachments also require separate inspection because text scanning cannot identify everything visible in an image.

Operationally, send only the context required for classification. Customer data and unpublished material remain subject to the application's data-sharing policy. Logs should retain an internal request reference, decision label, elapsed time, and error category rather than whole payloads or unfiltered error objects. A useful audit trail must remain useful without exposing credentials.

How should cost and latency be measured?

A token price is not a price per decision. At the time of this review, TypeSafe lists USD 42 per billion input tokens, equivalent to USD 0.042 per million input tokens. This is a provider price, not the project's observed bill, and it can change.

For a hypothetical calculation, 1,000 billable input tokens at that rate cost USD 0.000042 for that input portion. USD 0.000000042 is the price of one input token, not an arbitrary request. Confirm current billing rules and account for retries, fallback, System 2, infrastructure, and review effort when calculating completed-task cost.

A useful measurement set includes end-to-end p50 and p95 latency, request failures, fallback frequency, misclassification rates, and cost per accepted task. Compare alternatives on the same labeled requests under comparable conditions. A single sandbox call establishes neither reliability nor savings in production.

This update does not repeat old sandbox timing or savings figures as newly verified results. Without an appropriate measurement dataset, publish the measurement method and integration status instead. Adding more AI checks simply because the unit price is low can still increase total spending and waiting time. Efficiency should be judged at the workflow level.

When is this worth applying to your project?

Start with a recurring decision that is easy to label and low risk. Record the current baseline, run the alternative in shadow mode, examine disagreements, and only then consider expanding its authority. Define acceptance criteria before looking at the results so that favorable examples do not become the entire evaluation.

For actions affecting data or reputation, keep human review gates (Vietnamese) attached to the action itself. Correctly classifying a request does not mean its draft has been approved for publication. The central lesson from this review is to manage responsibility at each handoff: identify which engine ran, which outputs can be trusted, and which permissions remain under human control.

Frequently asked questions (FAQ)

Does System 1 replace System 2?

No. System 1 suits narrow decisions, while System 2 handles synthesis and planning. Deterministic rules still belong in code, and important actions need independent authorization controls.

Does having a Jev module mean deployment is complete?

No. Verify its call sites, shadow or active behavior, fallback path, and test results. Source integration alone does not establish that the same version is serving production or meeting performance targets.

Do typed outputs guarantee security?

No. Correct structure does not establish correct intent or permission. The application still needs authentication, data scoping, value validation, and a policy for missing results.

Can real logs be used as public examples?

Only after sanitization and a review of permission to disclose them. This article favors synthetic examples and excludes credential values, cookies, signatures, and customer identifiers from both language versions.

Lead Magnet Special Edition

Nhận Bộ Thư Viện Prompt & SOP AI Workflow Vận Hành Doanh Nghiệp 2026

Tặng miễn phí Ebook PDF + Notion Template quản lý AI System thực chiến từ Tôi Là Tùng. Gửi trực tiếp vào hòm thư công việc của bạn.

Bảo mật 100%• Nhận file PDF & Notion• Hủy đăng ký 1-Click
🎁 Miễn Phí & Trả Phí

Khám Phá Kho Workflow & SOP AI Thực Chiến

Thư viện quy trình n8n, Make.com và SOP vận hành AI tôi đang dùng thật — chọn đúng thứ bạn cần cho hệ thống của mình.

Nguyễn Thanh Tùng — AI System Designer
Written by Tùng
Nguyễn Thanh Tùng · AI Director