<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Evaluation on Lynx Tech Blog</title><link>https://blog.lynxflow.co/en/tags/evaluation/</link><description>Recent content in Evaluation on Lynx Tech Blog</description><generator>Hugo -- gohugo.io</generator><language>en</language><lastBuildDate>Sun, 23 Aug 2026 00:00:00 +0800</lastBuildDate><atom:link href="https://blog.lynxflow.co/en/tags/evaluation/index.xml" rel="self" type="application/rss+xml"/><item><title>Full Health Check of CPA Model Pool: Which of 80 Models Are Alive, Smart, or Running Naked</title><link>https://blog.lynxflow.co/en/posts/cpa-model-pool-audit-smart-model-ranking/</link><pubDate>Sun, 23 Aug 2026 00:00:00 +0800</pubDate><guid>https://blog.lynxflow.co/en/posts/cpa-model-pool-audit-smart-model-ranking/</guid><description>&lt;img src="https://blog.lynxflow.co/images/cpa-model-pool-audit-smart-model-ranking.png" alt="Featured image of post Full Health Check of CPA Model Pool: Which of 80 Models Are Alive, Smart, or Running Naked" /&gt;Last month I wrote an article testing six Claude Code models. This time I&amp;rsquo;m broadening the scope: my local CPA (local model gateway) has 80 models connected, and I did two things — first, sent real requests to each one to check liveness, then gave the survivors a puzzle-resistant intelligence test, and finally cross-validated against public leaderboards.
Bottom line up front: listing a model ≠ it works. Of the 80 models, 59 are text models, and only about 30 could actually hold a conversat</description></item><item><title>Hands-on Test of Six Claude Code Models: Who’s Fast, Who’s Smart, and Who’s a Lottery</title><link>https://blog.lynxflow.co/en/posts/claude-code-model-benchmark/</link><pubDate>Sat, 15 Aug 2026 00:00:00 +0800</pubDate><guid>https://blog.lynxflow.co/en/posts/claude-code-model-benchmark/</guid><description>&lt;img src="https://blog.lynxflow.co/images/cc-bench-hero.png?v=090603" alt="Featured image of post Hands-on Test of Six Claude Code Models: Who’s Fast, Who’s Smart, and Who’s a Lottery" /&gt;Claude Code’s model picker currently has six options sitting in it: qwen3.8-max, glm-5.2-fast-preview, glm-5.2, deepseek-v4-pro-0813, qwen3-coder-next, and grok-4.6. All of them are routed through my local CPA, a local model gateway, to their respective upstream providers. In day-to-day use, my impressions were vague: “this one feels faster,” “that one feels smarter.” Gut feel is unreliable, so I spent an evening putting them on the same starting line and benchmarked speed, thinking time, long-i</description></item></channel></rss>