Tại sao phần lớn agent workflow không cần intelligence 'frontier' cho mọi bước?
Hầu hết agentic system production (code agent, research agent, workflow automation) gồm chuỗi subtask: planning, retrieval, tool call formatting, data prep, summarization, verification, editing. Chỉ 20-40% bước thực sự cần 'suy nghĩ sâu' hoặc creativity cao cấp của model Mythos/Fable-class.
Phần còn lại là công việc có pattern, cần tốc độ và cấu trúc rõ: code infilling, markdown formatting, song song hóa sub-calls, extract structured data. Gọi frontier cho hết dẫn đến burn rate cao (như Fable 5 $10/$50/M + token consumption lớn) và latency không cần thiết.
DiffusionGemma đạt >1000 tokens/sec trên H100, 700+ trên RTX 5090, 500+ qua NVIDIA NIM miễn phí (xác nhận Simon Willison). 4× nhanh hơn autoregressive cùng regime, quantized <18GB VRAM, day-0 Unsloth/vLLM/HF.
DiffusionGemma và họ local fast model mang lại con số cụ thể gì?
DiffusionGemma (Google DeepMind, open weights Apache 2.0, 26B MoE A4B active) dùng discrete diffusion: sinh block ~256 tokens song song thay vì token-by-token. Kết quả thực tế từ developer:
1000 tokens/sec trên single H100
- 700+ tokens/sec trên RTX 5090
- 500+ tokens/sec xác nhận bởi Simon Willison qua NVIDIA NIM miễn phí
- ~4× nhanh hơn autoregressive model cùng regime (low batch, single user)
Quantized vừa 18GB VRAM, tích hợp sẵn vLLM/Unsloth/HF/Transformers, NVIDIA tối ưu day-0. Phù hợp nhất cho interactive local: inline editing trong IDE, parallel data prep, structured output nhanh.
Trade-off chất lượng và khi nào local 'đủ' / 'không đủ'
Trade-off chính: experimental, quality trailing autoregressive cùng size trên output dài, complex reasoning, hoặc creative open-ended. Nó 'tự sửa' tốt trong block nhưng đôi khi thiếu coherence toàn cục so với Fable/Opus.
Đủ dùng khi:
- Task có cấu trúc cao (JSON, code patch nhỏ, report template)
- Cần latency thấp hoặc throughput cao (real-time suggestion, batch sub-agent)
- Dữ liệu nhạy cảm cần giữ local
Không đủ khi:
- Long-horizon autonomous với ít guardrail (multi-file refactor lớn, incident root-cause phức tạp)
- Cần 'taste' hoặc judgment kinh nghiệm cao
- Output cuối cùng phải cực kỳ chính xác, ít room cho lỗi
- Latency đầu token <20-50ms (không network)
- Chi phí fixed sau capex/GPU rental
- Privacy & data sovereignty cao
- Quality đủ cho structured/edit/parallel
- Giới hạn bởi VRAM và experimental quality
- Chất lượng dẫn đầu long-horizon, complex
- Chi phí biến đổi, burn nhanh với reasoning
- Dễ scale, không quản lý infra
- Network latency + egress cost
- Dễ predict outcome, ít eval nội bộ
Chi phí thực tế: local GPU vs cloud token billing
Cloud: chi phí biến đổi theo token, càng 'thông minh' càng nghĩ nhiều → sinh nhiều token → hóa đơn tăng phi tuyến. Một agent harness phức tạp dễ fan-out thành chục triệu tokens.
Local: capex GPU (hoặc rental) cố định, sau đó marginal cost ~0. Nhưng phải quản lý: VRAM fit, quantization, serving stack, update model, observability.
Với DiffusionGemma + Kimi K2.7 Code (rẻ ~$1-4/M), 60-80% subtask có thể offload local hoặc specialist rẻ → tổng cost dễ giảm 3-5× so với all-frontier, đồng thời giảm p99 latency.
- 1DecomposePhân rã task thành sub-steps + estimate độ phức tạp
- 2Route localFast path: Diffusion/Kimi local cho prep, edit, parallel
- 3EscalateChỉ gọi frontier (Bedrock) khi cần depth hoặc judgment cao
- 4Verify & mergeTest, schema, human-in-loop cho critical step
Pattern hybrid routing thực tế cho Node/AWS stack
Triển khai đơn giản:
- Task decomposition + lightweight classifier (hoặc heuristic) phân loại: 'fast-local', 'specialist', 'frontier-reason'.
- Fast path: gọi DiffusionGemma local (hoặc self-host trên ECS/EC2 GPU spot) cho editing, prep, parallel subcalls.
- Specialist path: Kimi K2.7 Code hoặc tương đương cho coding-heavy.
- Escalate chỉ khi cần: Fable/Opus qua Bedrock với prompt cache 90%, hard USD cap per task, timeout.
- Merge + verification layer (unit test, schema check, human-in-loop cho critical).
Trên AWS: dùng Lambda/ECS cho router, EKS/EC2 GPU cho local serving (hoặc Bedrock cho frontier), EventBridge/SQS cho orchestration. Giữ state trong DynamoDB với idempotency.
DiffusionGemma experimental, trailing AR trên output dài/complex/creative. Dùng cho tốc độ và subtask có cấu trúc; giữ frontier cho final reasoning hoặc high-stakes decision. Luôn có fallback và verification layer.
Nên/không nên làm gì ngay?
Nên:
- Bắt đầu benchmark DiffusionGemma (qua NIM hoặc local) và Kimi K2.7 Code trên 2-3 subtask thực của agent hiện tại trong 1 tuần.
- Xây classifier/router đơn giản + spending cap trước khi scale agent.
- Đo lường: token saved, p95 latency, quality delta (pass rate, human review) trên workload thật.
Không nên:
- All-in local cho mọi thứ ngay từ đầu (chất lượng frontier vẫn dẫn trên hard tasks).
- Bỏ qua eval và observability khi introduce local path.
- Quên rằng hardware local cũng có 'cost' (depreciation, power, ops).
Kết luận: hybrid không phải trend — là cách duy nhất để agent production vừa thông minh khi cần, vừa bền vững về chi phí và latency. Bắt đầu từ subtask rõ ràng nhất, đo trước khi mở rộng.