Now
What I'm focused on right now. Updated when things change.
Last updated: August 20, 2026 · Beijing
Building
- —最近把主要工程精力放在 Agent Harness 和 Benchmark,持续研究 Context Router、Tool Registry、Permission Gate、Verifier、Trace,以及怎样让评测结果可复现。
- —md2wechat-skill 已达到 3,617 Stars。下一步准备适配 DeepSeek Harness,做一个 dsh-plugin,测试公众号发布流程怎样进入插件化 Agent Runtime。
- —发布 Weekly Wisereads,把 Readwise 每周高亮最多的文章、视频、PDF 和电子书做成中文深度解读,并跑通发现、研究、质量门禁和原子发布。
- —开始认真骑行和拍 Vlog,也在探索一款偏习惯、养成和反馈的 Agent-native 骑行 iOS App。希望让 AI 带来的效率,慢慢转成更好的生活节奏。
- —Most of my engineering attention is now on Agent Harnesses and benchmarks: context routers, tool registries, permission gates, verifiers, traces, and reproducible evaluation.
- —md2wechat-skill has reached 3,617 GitHub stars. The next experiment is a dsh-plugin for DeepSeek Harness, testing how the WeChat publishing workflow fits into a plugin-based Agent runtime.
- —Shipped Weekly Wisereads, turning Readwise's most-highlighted weekly articles, videos, PDFs, and curated ebooks into Chinese deep-reading reports with discovery, research, quality gates, and atomic publishing.
- —Started cycling seriously and filming vlogs. I am also exploring an Agent-native cycling iOS app centered on habits, progression, and feedback, while turning AI-enabled efficiency into a better pace of life.
Writing
Writing about Agent Harnesses, benchmarks, Skill engineering, open-source retrospectives, and Weekly Wisereads. Cycling is now part of the content plan too: first rides, city routes, and the question of where the time saved by AI actually goes.
Community
Continuing to run an AI Builder community around Agent Harnesses, benchmarks, open-source projects, md2wechat, Codex workflows, and lessons from building a real content growth system. Join my ZhiShiXingQiu: 杰尼·AI 实战圈. Join
Learning
- —继续拆解 DeepSeek Harness、Codex 和其他 Agent Runtime,重点看 Context、Memory、Tool、Permission、Verifier 和 Trace 怎样协同。
- —最近做了很多 Benchmark,正在研究怎样评估 Agent 找证据、用工具、修正错误和验证结论的完整过程。
- —把骑行当成一个长期实验,观察习惯、即时反馈和低摩擦操作怎样帮助人持续行动,也为 Agent-native 骑行产品积累真实输入。
- —Studying how context, memory, tools, permissions, verifiers, and traces work together across DeepSeek Harness, Codex, and other Agent runtimes.
- —After running many benchmarks, I am focusing on how to evaluate evidence gathering, tool use, error repair, and verification across the full Agent process.
- —Treating cycling as a long-term experiment in habits, immediate feedback, and low-friction interaction, while gathering real input for an Agent-native cycling product.
This is a now page, a concept by Derek Sivers.