<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>SWE-Bench on Lynx 的技术博客</title><link>https://blog.lynxflow.co/tags/swe-bench/</link><description>Recent content in SWE-Bench on Lynx 的技术博客</description><generator>Hugo -- gohugo.io</generator><language>zh-cn</language><lastBuildDate>Thu, 13 Aug 2026 00:00:00 +0800</lastBuildDate><atom:link href="https://blog.lynxflow.co/tags/swe-bench/index.xml" rel="self" type="application/rss+xml"/><item><title>选日常 Coding Agent 主力模型：别再只看 SWE-bench 第一</title><link>https://blog.lynxflow.co/posts/2026-08-13-coding-agent-benchmark-selection/</link><pubDate>Thu, 13 Aug 2026 00:00:00 +0800</pubDate><guid>https://blog.lynxflow.co/posts/2026-08-13-coding-agent-benchmark-selection/</guid><description>&lt;img src="https://blog.lynxflow.co/images/2026-08-13-coding-agent-benchmark-selection.png" alt="Featured image of post 选日常 Coding Agent 主力模型：别再只看 SWE-bench 第一" /&gt; 结论先讲：日常用 Claude Code / Codex 写代码，不该再把「SWE-bench 第一」当成选主力的唯一尺子。2026 年更靠谱的公开组合是 Terminal-Bench + SWE-rebench/DeepSWE + 成本看板；最终判决还得回到你自己的仓库任务。</description></item></channel></rss>