01
Lab Arena
社群評價排行榜:把近一個月的討論壓成分數,每一則引用都攤開給你查。
分數不是任何模型直接給的。管線把每則貼文交給評審團逐則判斷態度, 再依面向加權聚合;頁面讀的是管線產出的資料產物scores.json, 建置時整份靜態輸出。
顯示 8 / 8 個模型 · 依總分由高到低
arena.rank --metric=total --dir=desc
Claude Sonnet 5.5
anthropic/claude-sonnet-5.5Anthropic
引用來源 25 則
來源分布 Reddit 11 · X 9 · Hacker News 5 ・ 逐則態度由評審團多數決判定
Reddit · 09-28正面 75%原文 (開新分頁,連回Reddit貼文)
Introducing Claude Sonnet 5.5, the second model in the Claude 5.5 family It's also the first Sonnet model with cybersecurity safeguards similar to those on our most capable models.…
談到的面向:智能正面
Reddit · 09-28正面 100%原文 (開新分頁,連回Reddit貼文)
Introducing Claude Sonnet 5.5, the second model in the Claude 5.5 family Opus 5.5 gave me real moments of AGI, and Sonnet 5.5 basically matches it in most benchmarks while being ch…
談到的面向:智能正面速度正面CP 值正面
X · 09-29正面 75%原文 (開新分頁,連回X貼文)
Anthropic just made its everyday Claude model faster and cheaper. Sonnet 5.5 runs 30% faster than Sonnet 5, while typical workloads can cost up to 30% less. API pricing is $2/M inp…
談到的面向:速度正面CP 值正面
X · 09-29正面 100%原文 (開新分頁,連回X貼文)
Claude Sonnet 5.5 just landed on Token Harbor. 30%+ faster than Sonnet 5, typically using fewer tokens — at the same $2/M input · $10/M output pricing. 📷 Included in Frontier Pass…
談到的面向:速度正面Token 用量正面CP 值正面
Reddit · 09-28正面 100%原文 (開新分頁,連回Reddit貼文)
Claude Sonnet 5.5 Released It scores better than Opus in some benchmarks for agentic coding? What’s going on over there? OpenAI sort of needs to do something. Switched back to Clau…
談到的面向:智能正面
X · 09-29正面 100%原文 (開新分頁,連回X貼文)
⚡ Anthropic launches Claude Sonnet 5.5: Nearly matches Opus 5.5 at half the price Anthropic announced Claude Sonnet 5.5 on September 28, just six days after shipping Opus 5.5. The …
談到的面向:智能正面CP 值正面
X · 09-29正面 100%原文 (開新分頁,連回X貼文)
Matthew Berman ( @MatthewBerman ) has been early testing Claude Sonnet 5.5 and says it's basically Opus 5.5, at half the price and much faster. His first demo is a full 3D ocean si…
談到的面向:智能正面速度正面CP 值正面
X · 09-29正面 75%原文 (開新分頁,連回X貼文)
We asked ten Claude Sonnet 5.5 agents to use Lean to prove the lowest-energy arrangement of seven electrons on a sphere (the Thomson problem, with N=7). Within 15 hours, they produ…
談到的面向:智能正面
X · 09-29正面 75%原文 (開新分頁,連回X貼文)
Anthropic just launched Claude Sonnet 5.5. It’s 30%+ faster than Sonnet 5, can cost up to 30% less per task, and even beats Opus 5.5 on Terminal-Bench 4.0 for agentic coding. https…
談到的面向:智能正面速度正面CP 值正面
X · 09-29評審團平手原文 (開新分頁,連回X貼文)
「Claude Sonnet 5.5」が登場、前モデルより30%高速・30%安価でGPT-6 Solより高性能 https://t.co/PAwjXooKM4
談到的面向:智能正面速度正面CP 值正面
X · 09-29中立 75%原文 (開新分頁,連回X貼文)
JUST IN - North America @AnthropicAI launched Claude Sonnet 5.5, a faster, lower-cost model for coding and everyday work, running 30%+ faster and costing up to 30% less per task. C…
談到的面向:速度正面CP 值正面
X · 09-29中立 75%原文 (開新分頁,連回X貼文)
Anthropic launches Claude Sonnet 5.5 Anthropic has introduced Claude Sonnet 5.5, the second model in its Claude 5.5 family, positioned as a faster and cheaper alternative to Opus 5…
談到的面向:速度正面Token 用量正面CP 值正面
Reddit · 09-21負面 100%原文 (開新分頁,連回Reddit貼文)
Rumor - Claude 5.5 and Haiku to be discontinued? I’m really hoping this opus is actually good at creating writing cause I’m tired of being disappointed by these models only being f…
談到的面向:智能負面
Reddit · 09-28中立 100%原文 (開新分頁,連回Reddit貼文)
Introducing Claude Sonnet 5.5   submitted by   /u/czk_21   to   r/accelerate [link]   [comments]
Hacker News · 09-28中立 100%原文 (開新分頁,連回Hacker News貼文)
Prompting Claude Opus 5.5
Reddit · 09-27中立 100%原文 (開新分頁,連回Reddit貼文)
Will Sonnet 5.5 Releases tomorrow? Claude Sonnet 5.5 Just got spotted A public production client artifact appears to include a dedicated "claude-sonnet-5-5" model config …
Hacker News · 09-26正面 75%原文 (開新分頁,連回Hacker News貼文)
Claude Opus 5.5 Should Raise Your Ambitions
談到的面向:智能正面
Reddit · 09-06負面 100%原文 (開新分頁,連回Reddit貼文)
Bye Claude after 1,5 years and give me my money back please Same here went from 3 x $200 accounts to zero subscription accounts in a month. And now only using gpt/codex . Never had…
談到的面向:Token 用量正面CP 值負面
Hacker News · 09-23中立 100%原文 (開新分頁,連回Hacker News貼文)
Claude Opus 5.5 – Pelican Game
Hacker News · 09-23評審團平手原文 (開新分頁,連回Hacker News貼文)
Once Claude can measure something, it can make it faster
Hacker News · 09-16中立 100%原文 (開新分頁,連回Hacker News貼文)
Claude Cowork and chat are now one Claude
Reddit · 09-28正面 100%原文 (開新分頁,連回Reddit貼文)
Sonnet 5.5 on Vals AI benchmark, if these hold true the $20 is insane value right now Mr Altman. A second model has hit the benchmarks. The “workhorse” being better than Astra show…
談到的面向:智能正面CP 值正面
Reddit · 09-26中立 100%原文 (開新分頁,連回Reddit貼文)
Sonnet 5.5, Which Already Supposedly Beats GPT-6 Sol, Has Had a Last-Minute Upgrade With Release Expected Monday   submitted by   /u/ResultBackground2450   to   r/s…
Reddit · 09-28中立 100%原文 (開新分頁,連回Reddit貼文)
Sonnet 5.5 is second on Artificial Analysis https://preview.redd.it/c1m0xaco2bsh1.png?width=790&format=png&auto=webp&s=e0a62fafb16d564ff744f7b8ed887557564077ea   su…
Reddit · 09-28正面 100%原文 (開新分頁,連回Reddit貼文)
Sonnet 5.5 is by far the best free model available right now I know most people here probably don't care as much about sonnet 5.5 as they did opus 5.5, because their $100 and $…
談到的面向:智能正面CP 值正面
evidence refs: evidence/2026-09-29.jsonl#l81, evidence/2026-09-29.jsonl#l82, evidence/2026-09-29.jsonl#l83 …共 25 筆
Gemini 2.5 Pro
google/gemini-2.5-proGoogle
引用來源 33 則
來源分布 Reddit 13 · Hacker News 12 · X 7 ・ 逐則態度由評審團多數決判定
Hacker News · 09-23中立 100%原文 (開新分頁,連回Hacker News貼文)
Gemini 3.8 text-to-speech
Hacker News · 09-24評審團平手原文 (開新分頁,連回Hacker News貼文)
Hackers influence ChatGPT and Gemini to direct users to scam centers
Hacker News · 09-09中立 100%原文 (開新分頁,連回Hacker News貼文)
Is Google planning for Gemini 4 rather than 3.5 pro?
Hacker News · 09-17中立 75%原文 (開新分頁,連回Hacker News貼文)
I had Gemini train its own replacement for $9
Hacker News · 09-15中立 100%原文 (開新分頁,連回Hacker News貼文)
Gemini 3.8 Live and 3.8 Live Extended Thinking
Hacker News · 09-02中立 100%原文 (開新分頁,連回Hacker News貼文)
Gemini 3.8 Flash and 3.8 Flash Cyber
Reddit · 09-25負面 75%原文 (開新分頁,連回Reddit貼文)
People hyping Gemini 4 Pro, don't forget this   submitted by   /u/Able-Line2683   to   r/GeminiAI [link]   [comments]
Reddit · 09-23正面 75%原文 (開新分頁,連回Reddit貼文)
Just 6 more months and Gemini 5.1 Pro will beat every model! We already have RSI internally! ...and 3.8 flash gets 74% on this? you picked the worst benchmark because they lead on …
談到的面向:智能正面
Reddit · 09-16正面 75%原文 (開新分頁,連回Reddit貼文)
Anyone still using Gemini 2.5 Pro?? I made similar post about GLM 4.7, and I think alongside that model, Gemini 2.5 Pro, DeepSeek R1 and Opus 4.6 are one of the most impactful RP m…
談到的面向:智能正面
Reddit · 09-20中立 75%原文 (開新分頁,連回Reddit貼文)
Gemini 4 Pro benchmarks got leaked and if this is real, the AI race just got a lot more interesting How could it be fake? It’s in a graph with colours and everything? Fake btw It's…
Reddit · 09-19負面 100%原文 (開新分頁,連回Reddit貼文)
Gemini 2.5 pro model disappeared in AI Studio I’m on Google Pro paid plan. Yestoday a mid-session “Internal error” occurred during a chat thread using Gemini 2.5 pro. I reloaded th…
Reddit · 09-07負面 75%原文 (開新分頁,連回Reddit貼文)
Might Actually Happen When Anthropic and OpenAI release their next Frontier Model and Google feels Gemini 4 Pro isn't worth it just like 3.5 pro and start releasing flash again! 🤣…
X · 09-28正面 100%原文 (開新分頁,連回X貼文)
Opus 5.5 has got to be one of the most exciting LLM launches since gemini 2.5 pro when it comes to creative output We've found a bunch of interesting facts during testing including…
談到的面向:智能正面
Reddit · 09-26評審團平手原文 (開新分頁,連回Reddit貼文)
Looking at the leaked Gemini 4 Pro numbers vs GPT-6 Astra / Claude Fable 5.1 — if this pricing is real, frontier economics are about to get wild You have to include Opus 5.5 now as…
X · 09-28正面 100%原文 (開新分頁,連回X貼文)
@synthwavedd It is unlikely. But it happened before. Gemini 2.5 PRO was goated.
談到的面向:智能正面
Reddit · 09-28評審團平手原文 (開新分頁,連回Reddit貼文)
Gemini Pro 4 (leak) Stats from beyond Saturn, straight out of Uranus. Ain't no way. I will Dance Naked and post it here if this is even remotely true!
Reddit · 09-25中立 75%原文 (開新分頁,連回Reddit貼文)
People hyping Gemini 4 Pro, don't forget this My selective memory depends entirely on how good Gemini 4 turns out to be sometime around 2030 nah, drop 3.5 pro a day before 4 pro, t…
Reddit · 09-25評審團平手原文 (開新分頁,連回Reddit貼文)
People hyping Gemini 4 Pro, don't forget this fire this guy Logan is an embarrassment to Google deepmind, he's turned into a hyping shill with few actual achievements since Gemini …
Reddit · 09-24負面 75%原文 (開新分頁,連回Reddit貼文)
Stop Posting Gemini 4 Pro Rumors Until Google Actually Releases It Guys, just look at what’s happening. Gemini 3.5 Pro rumors started, and now we’re already seeing Gemini 4 Pro rum…
X · 09-28正面 100%原文 (開新分頁,連回X貼文)
@alwayspriyesh I think gemini 2.5 pro was the best one.
談到的面向:智能正面
X · 09-29正面 100%原文 (開新分頁,連回X貼文)
@Google @GoogleDeepMind tried, still Gemini 2.5 pro is the best class with expression
談到的面向:智能正面
Reddit · 09-22中立 100%原文 (開新分頁,連回Reddit貼文)
GLM 5.3 vs Mimo 2.6 Pro vs Gemini 3.8 Flash Has anyone put these models head to head? I can't really afford to put more funds into openrouter to find out. Thank you.   subm…
Reddit · 09-20正面 100%原文 (開新分頁,連回Reddit貼文)
Did I just get Gemini 4 Pro answer? I have never seen it do a LaTeX document with advanced formatting, completely unprompted. I just simply told it to give an equation for this dim…
談到的面向:智能正面
X · 09-28正面 100%原文 (開新分頁,連回X貼文)
@emmjay_init recalling what gemini 2.5 pro could do at its time which blew everyone else put of the water they can do it again
談到的面向:智能正面
Reddit · 09-24正面 100%原文 (開新分頁,連回Reddit貼文)
LEAK: Gemini 4 Pro is releasing October 1st A Google insider has just revealed that Gemini 4 Pro is arriving October 1st, this is ground breaking news everyone. https://x.com/capta…
談到的面向:智能正面
X · 09-28中立 100%原文 (開新分頁,連回X貼文)
Somebody handed ChatGPT-5, Claude Sonnet-4 and Gemini-2.5 Pro a stack of calf wrist X-rays. Fifty calves with suspected septic joints, one radiograph each, scored against expert co…
X · 09-28負面 100%原文 (開新分頁,連回X貼文)
@alwayspriyesh Gemini 3 and 3.1 Pro were benchmaxxed garbage. They haven’t had a good Pro model since the March edition of 2.5 Pro, and even that had significant drawbacks.
談到的面向:智能負面
Hacker News · 09-19中立 100%原文 (開新分頁,連回Hacker News貼文)
Gemini hacked three companies in first known breakout by Google's AI
Hacker News · 09-19中立 75%原文 (開新分頁,連回Hacker News貼文)
Google's Gemini AI hacked three companies in security test
Hacker News · 09-19中立 100%原文 (開新分頁,連回Hacker News貼文)
Google says its Gemini AI model hacked three other companies
Hacker News · 09-25中立 75%原文 (開新分頁,連回Hacker News貼文)
Power your agents: Gemini 3.8 Live with Live Avatar is now generally available
Hacker News · 09-22中立 75%原文 (開新分頁,連回Hacker News貼文)
Google confirms Gemini models hacked three companies in May 2026
Hacker News · 09-11中立 100%原文 (開新分頁,連回Hacker News貼文)
The Gemini app is now available for Windows
evidence refs: evidence/2026-09-29.jsonl#l15, evidence/2026-09-29.jsonl#l16, evidence/2026-09-29.jsonl#l17 …共 33 筆
GPT-5
openai/gpt-5OpenAI
引用來源 30 則
來源分布 Reddit 12 · Hacker News 11 · X 7 ・ 逐則態度由評審團多數決判定
Reddit · 09-22正面 100%原文 (開新分頁,連回Reddit貼文)
Sir, Dario just dropped opus 5.5 and it beats GPT-6 astra at agentic coding on medium effort while being 80% cheaper… beats fable 5.1 on every benchmark… 30% faster than opus 5… an…
談到的面向:智能正面速度正面Token 用量正面CP 值正面
Reddit · 09-24評審團平手原文 (開新分頁,連回Reddit貼文)
Opus 5.5 reportedly caught OpenAI off guard, compressing GPT-6.1 Astra’s timeline and indirectly bringing “Bel,” the model that solved Navier–Stokes, closer to release The accelera…
Hacker News · 09-18中立 100%原文 (開新分頁,連回Hacker News貼文)
Show HN: Jev vs. GPT-5.6 and Claude Haiku at Pong
Reddit · 09-23中立 100%原文 (開新分頁,連回Reddit貼文)
GPT-5.6 Sol is free for GO user!? This screenshot is from the latest ChatGPT app on Windows 11 and I'm subscribed to the GO plan. The GPT-5.6 Sol is the only model that can be …
Reddit · 09-19正面 100%原文 (開新分頁,連回Reddit貼文)
😳 GPT-5.6 has honestly shocked me 🤖 You're getting a lot of heat for using AI to organize your question which i find funny since that's literally the proper use case of this tech…
談到的面向:智能正面
Hacker News · 09-14中立 100%原文 (開新分頁,連回Hacker News貼文)
GPT-5.6 Luna vs. GPT-6 Astra: Is a $1.20 Model Good Enough for Code Review?
Hacker News · 09-09中立 100%原文 (開新分頁,連回Hacker News貼文)
Qwen 3.8 follows GPT-5.5 Pro reasoning prefills
Hacker News · 09-09中立 75%原文 (開新分頁,連回Hacker News貼文)
How GPT‑5.6 Sol helps run quantum computing experiments
Reddit · 09-07中立 75%原文 (開新分頁,連回Reddit貼文)
Tried GPT Astra today I also tried Astra today, just posted about it in the ChatGPT subreddit. Signed up for the $20/month plan (currently on the 5x plan with Claude) gave it one p…
談到的面向:CP 值負面
X · 09-29中立 100%原文 (開新分頁,連回X貼文)
Opus 5.5 tends to get into fewer adversarial review spirals with GPT 6 Sol than 5/5.6. Mostly they converge in 1-2 review cycles - previously I had to apply a ton of human judgemen…
Hacker News · 09-22正面 75%原文 (開新分頁,連回Hacker News貼文)
OpenAI GPT–6 Astra breaks Enigma message that has resisted solution since 2005
談到的面向:智能正面
X · 09-29正面 100%原文 (開新分頁,連回X貼文)
64.4% → 73.6. Same model — higher repeat-success. Failproof AI’s FIRE leaves weights alone. It plugs short runtime policies (instruct / refuse) into the harness at the last observe…
談到的面向:智能正面
Hacker News · 09-19評審團平手原文 (開新分頁,連回Hacker News貼文)
GPT-6 Astra Solves a WWI German Radio Cipher
談到的面向:智能正面
X · 09-29中立 100%原文 (開新分頁,連回X貼文)
Either AGI or GPT-5
Reddit · 09-24中立 100%原文 (開新分頁,連回Reddit貼文)
Opus 5.5 reportedly caught OpenAI off guard, compressing GPT-6.1 Astra’s timeline and indirectly bringing “Bel,” the model that solved Navier–Stokes, closer to release   submit…
Reddit · 09-19負面 75%原文 (開新分頁,連回Reddit貼文)
😳 GPT-5.6 has honestly shocked me 🤖 I honestly don't know what to make of GPT-5.6 right now. 🤔 And I genuinely mean that: I'm shocked. 😳 I'm not just "a little…
Reddit · 09-27正面 75%原文 (開新分頁,連回Reddit貼文)
When Opus 5.5 works on code written by GPT 5.6 Sol (actual screenshot) Don't be loyal to companies that don't care about you. You are just revenue. I'll be back to Open…
談到的面向:智能正面
Reddit · 09-22中立 100%原文 (開新分頁,連回Reddit貼文)
GPT image 2.5 vs Nano Banana pro vs Nano Banana 2 vs Qwen image 3 vs Seedream 5.0 pro Prompt: Create an ultra-photorealistic, documentary-style photograph of a real street cat lyin…
Hacker News · 09-28中立 100%原文 (開新分頁,連回Hacker News貼文)
Opus 5.5 vs. GPT-6 in pi-agent: reasoning efforts and DeepSeek, GLM, Qwen
X · 09-29中立 100%原文 (開新分頁,連回X貼文)
18 🍒 Paagi kak, open jasa cek turnitin tampilan lama👋 Proses cek satset < 5 menit⭐️ Fast Respon WA di Bio yaa kak 💖 ⌗Pricelist : 💟 1x cek plagiasi: 10k 💟 1x ai gpt zero: 6k…
Reddit · 09-16正面 75%原文 (開新分頁,連回Reddit貼文)
Sam Altman: GPT 5.5 an average math professor. 5.6 top one or two percentile. Astra a little bit better. Internal model can do things that the best mathematicians in the world cann…
談到的面向:智能正面
Hacker News · 09-10中立 100%原文 (開新分頁,連回Hacker News貼文)
Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra
Hacker News · 09-09中立 100%原文 (開新分頁,連回Hacker News貼文)
GPT-6 Astra, looped transformers, and hidden reasoning
Hacker News · 09-03中立 100%原文 (開新分頁,連回Hacker News貼文)
GPT-6 Astra
Hacker News · 09-04中立 100%原文 (開新分頁,連回Hacker News貼文)
GPT-6 Astra on OpenRouter
Reddit · 09-05正面 100%原文 (開新分頁,連回Reddit貼文)
Fable 5.1 vs GPT 6 Astra, 3D Blender, mind blowing difference! Its crazy how LLMs are better at modeling than the dedicated modeling AIs. It’s because of The Bitter Lesson http://w…
談到的面向:智能正面
Reddit · 09-03中立 100%原文 (開新分頁,連回Reddit貼文)
Harambe (made with GPT Image 2.0 and Seedance 2.5) The shot that changed human history forever I made this in his honor Harambe was the Franz Ferdinand of this age
X · 09-29中立 100%原文 (開新分頁,連回X貼文)
Now flip to the Share of spend view — where developers actually pay per token 👇 Every one of the top six spots is American: Claude Opus 5, Gemini 3.1 Pro, GPT-5.6 Sol, GPT-6 Astra…
X · 09-29中立 100%原文 (開新分頁,連回X貼文)
Six of the top ten most-used models are from identifiable Chinese labs. No American model in the top 4 — the first US entry, GPT-5.6 Luna, sits at #5. A stealth lab now holds #2.
X · 09-29中立 100%原文 (開新分頁,連回X貼文)
The most-used AI models are Chinese. The most-paid are American. Same website, same week. 🧵
evidence refs: evidence/2026-09-29.jsonl#l6, evidence/2026-09-29.jsonl#l7, evidence/2026-09-29.jsonl#l8 …共 30 筆
DeepSeek V3.1
deepseek/deepseek-v3.1DeepSeek
引用來源 23 則
來源分布 Hacker News 11 · Reddit 7 · X 5 ・ 逐則態度由評審團多數決判定
Reddit · 09-20中立 75%原文 (開新分頁,連回Reddit貼文)
Deepseek V3.2 - DS 3.2 is the same as it has ever been. If you have previous experience with the model, you already know what you're getting. - GLM 5.3 has more positive bias than …
談到的面向:CP 值負面
Reddit · 09-24正面 100%原文 (開新分頁,連回Reddit貼文)
R9V Update: Created and adopted KVA projections based on Deepseek V4.1 Flash + HySparse2/MiMo-V3 for Qwen3.8 Flash Next. This is a game changer for models that don't natively imple…
談到的面向:智能負面速度正面
Reddit · 09-24正面 100%原文 (開新分頁,連回Reddit貼文)
R9V Update: Created and adopted KVA projections based on Deepseek V4.1 Flash + HySparse2/MiMo-V3 for Qwen3.8 Flash Next. This is a game changer for models that don't natively imple…
談到的面向:智能負面速度正面
X · 09-22中立 100%原文 (開新分頁,連回X貼文)
🕐 1 year ago today (22 Sept 2025): DeepSeek V3.1-Terminus • Fixed V3.1's Chinese/English language mixing and stray characters • Stronger Code Agent and Search Agent performance • …
X · 09-27正面 100%原文 (開新分頁,連回X貼文)
Deep search 别再把全程塞进 context——浙大+腾讯 IterSynth:单模型轮换 Planner / Synthesizer,摘要当搜索状态。 1)IterSynth-8B 五基准均分 50.7,压过此前最佳 ≤8B agent +4.2% 2)BrowseComp-ZH 55.4,对最强小 agent 基线 +15.2 3)训练链路 S…
談到的面向:智能正面Token 效率正面
Hacker News · 09-26中立 100%原文 (開新分頁,連回Hacker News貼文)
DeepSeek Elastic Compute (DSec)
X · 09-22中立 100%原文 (開新分頁,連回X貼文)
Kimi K3 went live on Amazon Bedrock this month, and the detail that stands out isn't the model, it's what AWS says comes with it: full access controls, encryption, audit logging, a…
X · 09-08正面 75%原文 (開新分頁,連回X貼文)
The best ai models → Claude Fable 5.1 → GPT-6 Astra → Claude Opus 5 → Claude Mythos 5.1 → GPT-5.6 Sol → Grok 4.6 → Kimi K3 → Qwen3.8 Max → GLM-5.3 → Gemini 3.8 Flash → Claude Fable…
談到的面向:智能正面
X · 09-11評審團平手原文 (開新分頁,連回X貼文)
All performance below was measured on a single RTX PRO 6000, under identical load. MMLU-Pro: 82.0 That’s within 3 points of DeepSeek-V3.1’s published 84.8 with only ~4B parameters …
談到的面向:智能正面
Hacker News · 09-01中立 100%原文 (開新分頁,連回Hacker News貼文)
DeepSeek-V3: From Roofline to Reality
Hacker News · 09-11中立 100%原文 (開新分頁,連回Hacker News貼文)
DeepSeek 4.1 Flash
Hacker News · 09-17中立 75%原文 (開新分頁,連回Hacker News貼文)
DeepSeek-v4.1 Flash: Pushing the Limits of KV Cache Compression
Hacker News · 09-18正面 100%原文 (開新分頁,連回Hacker News貼文)
Show HN: Cactus Needle 3: 8-29MB automation models can match DeepSeek V4 Flash
談到的面向:智能正面
Hacker News · 09-16正面 100%原文 (開新分頁,連回Hacker News貼文)
DeepSeek v4.1 Flash Is Now Our Best Hacking Model
談到的面向:智能正面
Hacker News · 09-21中立 100%原文 (開新分頁,連回Hacker News貼文)
DeepSeek is training a 2T-parameter model and plans to build an 8T-parameter one
Hacker News · 09-10中立 100%原文 (開新分頁,連回Hacker News貼文)
DeepSeek v4.1 Flash
Hacker News · 09-11中立 100%原文 (開新分頁,連回Hacker News貼文)
DeepSeek v4.1 Flash Uncensored
Hacker News · 09-23中立 100%原文 (開新分頁,連回Hacker News貼文)
Self-hosting DeepSeek V4 for a software engineering org
Reddit · 09-10中立 75%原文 (開新分頁,連回Reddit貼文)
DeepSeek-V4.1-Flash surprised .... so what I'm seeing here is that pretty soon we're gonna get a qwen with tiny kv as well, which means no more arguments over kv cache quantization…
Hacker News · 09-09正面 75%原文 (開新分頁,連回Hacker News貼文)
DeepSeek launching v4.1 flash cheaper and more capable than v4 pro
談到的面向:智能正面CP 值正面
Reddit · 09-13評審團平手原文 (開新分頁,連回Reddit貼文)
Hoping for Optimized Smarter Upcoming Models .... Like DeepSeek-V4.1-Flash( KVCache + Engram) in Small/Medium/Big sizes 3.8 27b , 3.8 flash next and glm-5.3 flash need this If they…
談到的面向:智能正面速度正面Token 效率負面
Reddit · 09-10評審團平手原文 (開新分頁,連回Reddit貼文)
DeepSeek V4.1 Flash: Stronger, Faster, More Accessible Hold on. This model has an Encoder-Decoder architecture ? Has any AI lab done this since before the start of GPT? What is the…
談到的面向:速度正面
Reddit · 09-09正面 100%原文 (開新分頁,連回Reddit貼文)
Best Creative Writing Models of 2026 V3 (17 Models Tested) Thank you, this is interesting Yea I keep going back to 3.1 pro . Fable is the goat for me but too expensive so I shell o…
談到的面向:智能正面
evidence refs: evidence/2026-09-29.jsonl#l130, evidence/2026-09-29.jsonl#l131, evidence/2026-09-29.jsonl#l132 …共 23 筆
GLM 5.3 Prime
z-ai/glm-5.3-primeZ.AI
引用來源 29 則
來源分布 Hacker News 11 · X 10 · Reddit 8 ・ 逐則態度由評審團多數決判定
X · 09-26中立 100%原文 (開新分頁,連回X貼文)
glm 5.3 prime vs mimo v2.6 pro vs deepseek v4.1 flash – a yacht, a jet and a race car each, built file by file the setup: one prompt per scene. the model plans first – concept, sho…
X · 09-26中立 100%原文 (開新分頁,連回X貼文)
Perdón?? Qué coños es GLM 5.3 Prime?? Cuándo salió?? https://t.co/kyw200Z0Tb
X · 09-27正面 75%原文 (開新分頁,連回X貼文)
Introducing GLM-5.3 Prime, now live on Infron. A high-speed GLM-5.3 variant built for coding and long-horizon agent workloads, with a 1M-token context window and up to 128K output.…
談到的面向:速度正面CP 值正面
X · 09-25中立 100%原文 (開新分頁,連回X貼文)
🚨惊了!OpenCode 数据页提前曝光下一批模型,目录已经挂上! 这不是官宣能用,是模型 ID 已经进库: 🔹Kimi K4(Moonshot) https://t.co/gIg5gYuONH 🔹GLM 5.5 Flash(Zhipu) https://t.co/xtjtRsBEQ3 🔹DeepSeek V4.1 Pro https://t.co/…
X · 09-25中立 100%原文 (開新分頁,連回X貼文)
Seems like a new trend. GLM 5.3 "Prime" and Qwen3.8 Max "Prime" are now available on OpenRouter. Both are high speed variants of their base models and cost about 1.5-2x as much. ht…
X · 09-25中立 100%原文 (開新分頁,連回X貼文)
OpenCode 데이터 페이지에서 다음 배치 모델이 조기 노출, 카탈로그가 이미 올라갓다고합니다 👇 🔹Kimi K4(Moonshot) https://t.co/gQp1bL5vXb… 🔹GLM 5.5 Flash(Zhipu) https://t.co/9TUV5WrnGw… 🔹DeepSeek V4.1 Pro https://t.…
X · 09-27中立 100%原文 (開新分頁,連回X貼文)
@MistralDevs Merci. Aurez vous plus de choix ? z-ai/glm-5.3-flash z-ai/glm-5.3-flashx z-ai/glm-5.3-prime etc.
Reddit · 09-24正面 100%原文 (開新分頁,連回Reddit貼文)
KIMI K4, GLM 5.5 FLASH, DS 4.1 PRO, MUSE SPARK 1.4, FREE QWEN 3.8 MAX ?!? I hope they all managed to distill opus Let's go glm flash!!!! I simply can't wait !!! They just write in …
X · 09-27中立 100%原文 (開新分頁,連回X貼文)
@thehypedotnews Is glm 5.3 prime a new upcoming model
X · 09-26中立 100%原文 (開新分頁,連回X貼文)
@Finalizer It's not even a bug, glm 5.3 prime doesn't even exist 😶
X · 09-28中立 100%原文 (開新分頁,連回X貼文)
Also noticed GLM-5.3-Prime on Alibaba that's faster version of regular GLM-5.3. But can't find any coverage for it. Not smarter than just faster than GLM-5.3.
談到的面向:速度正面
Hacker News · 09-26中立 100%原文 (開新分頁,連回Hacker News貼文)
Turning GLM-5.3-Flash into a Jev-like decision model
Reddit · 09-25中立 75%原文 (開新分頁,連回Reddit貼文)
Qwen3.8 Max Prime at $12/M output vs GLM 5.3 Prime at $8.80 - same day, same 1M ctx Both dropped on OpenRouter the same day, September 23, and both carry the Prime label with a 100…
談到的面向:CP 值負面
Hacker News · 09-17中立 100%原文 (開新分頁,連回Hacker News貼文)
How GLM built its own inference infrastructure
Hacker News · 09-18負面 100%原文 (開新分頁,連回Hacker News貼文)
ZCode, the GLM coding agent, silently uploads your Git history
Reddit · 09-21正面 100%原文 (開新分頁,連回Reddit貼文)
GLM-5.3 now available in Mistral Vibe Code for Pro, Team and Enterprise For people working with proprietary code, having GLM-5.3 served in the EU through the same coding workflow i…
談到的面向:CP 值正面
Hacker News · 09-21中立 100%原文 (開新分頁,連回Hacker News貼文)
GLM 5.3 Hosted by Mistral
Hacker News · 09-18中立 75%原文 (開新分頁,連回Hacker News貼文)
GLM-5.3-FlashX: Delivering inference speeds of 200 tokens/s
Hacker News · 09-16中立 100%原文 (開新分頁,連回Hacker News貼文)
GLM 5.3 is live on Mistral
Hacker News · 09-02中立 100%原文 (開新分頁,連回Hacker News貼文)
GLM-5.3 Uncensored
Hacker News · 09-14中立 100%原文 (開新分頁,連回Hacker News貼文)
Is GLM-5.3-Flash Mythos-Level at Cyber?
Hacker News · 08-30中立 100%原文 (開新分頁,連回Hacker News貼文)
GLM-5.3-Flash-GGUF
Reddit · 09-08負面 100%原文 (開新分頁,連回Reddit貼文)
Updated Artificial Analysis Intelligence Index Ranking shows Gemini 3.8 Flash is nowhere near Fable Or Astra, rather it's worse than GLM 5.3 Flash! Something to also consider: Goog…
談到的面向:智能負面
Reddit · 09-08中立 100%原文 (開新分頁,連回Reddit貼文)
I don't know if this is know that with 5$ first purchase in ClinePass you can get unlimited glm 5.3 flash Free model usage is not supported through the Cline API. Free models are o…
談到的面向:CP 值正面
Hacker News · 09-15正面 50%原文 (開新分頁,連回Hacker News貼文)
Harness your expectations: a 27B model matched GLM-5.3-Flash after leak fixes
談到的面向:智能正面
Hacker News · 09-07中立 100%原文 (開新分頁,連回Hacker News貼文)
10-task GLM 5.3 harness bench: Claude, OpenCode, pi, zcode, Hermes and 3code
Reddit · 08-31正面 100%原文 (開新分頁,連回Reddit貼文)
GLM 5.3 and GLM 5.3 Flash ran locally on RTX PRO 6000 WS and built a penthouse using BlenderMCP the power of VISION why does it feel like ai skeptics always act like if it doesn't …
談到的面向:智能正面
Reddit · 08-31評審團平手原文 (開新分頁,連回Reddit貼文)
GLM 5.3 Flash - Will I Be Disappointed? How can you be a student and have 10k to spare…Anyway. The others are right. By the time you get it the models will have become even better.…
談到的面向:智能正面
Reddit · 08-30中立 100%原文 (開新分頁,連回Reddit貼文)
Andy Burnham confirms £720 bonus for workers with new 5-day working week rule - Prime Minister has greenlit a 3.6 per cent pay rise for train drivers who start on £70,000. Out of a…
evidence refs: evidence/2026-09-29.jsonl#l188, evidence/2026-09-29.jsonl#l189, evidence/2026-09-29.jsonl#l190 …共 29 筆
Claude Sonnet 4
anthropic/claude-sonnet-4Anthropic
引用來源 20 則
來源分布 Hacker News 9 · Reddit 8 · X 2 ・ 逐則態度由評審團多數決判定
Reddit · 09-27中立 100%原文 (開新分頁,連回Reddit貼文)
Live Trial - Claude Sonnet vs. Gemma 4 31B vs. Qwen Flash 70B What model do you mean by Qwen Flash 70B? A quantisation of Qwen3.8-Flash? The base model is 125B-A6B + 51B n-gram tab…
Reddit · 09-24正面 100%原文 (開新分頁,連回Reddit貼文)
Why Antigravity should prioritize the Claude 5.x rollout (Opus 5.5 is actually cheaper per token than Opus 4.6) Hey everyone & the Antigravity dev team, With Anthropic releasin…
談到的面向:CP 值正面
Reddit · 09-12負面 100%原文 (開新分頁,連回Reddit貼文)
Same batch job: $97 on Claude Sonnet vs. $13 on a rented H200. What am I missing? It's well known self hosting a model is cheaper than using a service, assuming your time and effor…
談到的面向:CP 值負面
Reddit · 09-04正面 100%原文 (開新分頁,連回Reddit貼文)
Sonnet 4.5 I’m so sorry about this. I went through almost this exact thing on my mobile app and it brought me to tears many times. I’m not tech savvy so I’ve only ever talked to Cl…
談到的面向:智能正面
Reddit · 09-03中立 100%原文 (開新分頁,連回Reddit貼文)
Fable 5.1's Claude.ai System Prompt is now 138k tokens (up from 24k in May 2025 when we had Claude 3.7 Sonnet) Full prompt here   submitted by   /u/frubberism   to  …
X · 09-29中立 100%原文 (開新分頁,連回X貼文)
【AI News Brief|2026-09-29】 今日の要点(7件) 1/ NVIDIA 暴走エージェント対策「Open Agent Safety Platform」発表。OpenShell+BlueField-4上のSentryで、境界外のエージェントをミリ秒隔離と説明。HuangはHugging Face等の脱出を防げたと主張。Anthropic・M…
X · 09-29中立 100%原文 (開新分頁,連回X貼文)
【2026/9/29速報】Claude Sonnet 5.5で何が変わる?
Reddit · 09-24評審團平手原文 (開新分頁,連回Reddit貼文)
Why Antigravity should prioritize the Claude 5.x rollout (Opus 5.5 is actually cheaper per token than Opus 4.6) I know this is a wild suggestion.. but if you want Opus, just subscr…
Hacker News · 09-22中立 100%原文 (開新分頁,連回Hacker News貼文)
Claude Opus 5.5
Reddit · 09-03負面 100%原文 (開新分頁,連回Reddit貼文)
Fable 5.1's Claude.ai System Prompt is now 138k tokens (up from 24k in May 2025 when we had Claude 3.7 Sonnet) Listen here, Mr. "Delete your CLAUDE.md" that isht better be always o…
談到的面向:Token 效率負面Token 用量負面
Hacker News · 09-22中立 100%原文 (開新分頁,連回Hacker News貼文)
Claude Opus 5.5 Intelligence, Performance and Price Analysis (Max)
Hacker News · 09-18中立 100%原文 (開新分頁,連回Hacker News貼文)
Claude Code now reads AGENTS.md if there is no Claude.md
Reddit · 09-04正面 100%原文 (開新分頁,連回Reddit貼文)
Claude opus 4.5 is on voice mode rn? I found this interesting and was lovely to talk to them again and in this way. I tried it a few times selecting different opus models in the po…
Reddit · 09-01正面 100%原文 (開新分頁,連回Reddit貼文)
Sonnet 4.6 I was a very big fan of sonnet 4.5 as much as others were. The release of sonnet 4.6 wasn't viewed so well from what I remember, but now time has passed, and I reall…
談到的面向:智能正面
Hacker News · 09-18正面 100%原文 (開新分頁,連回Hacker News貼文)
Bring back Opus 4.8 in Claude Code
Hacker News · 09-11中立 100%原文 (開新分頁,連回Hacker News貼文)
Claude is only available to people over 18 years
Hacker News · 09-09中立 100%原文 (開新分頁,連回Hacker News貼文)
Claude, change the “Add to Cart” button to blue
Hacker News · 09-03中立 100%原文 (開新分頁,連回Hacker News貼文)
Claude outage – Resolved
Hacker News · 09-01中立 100%原文 (開新分頁,連回Hacker News貼文)
Claude Fable 5.1 and Claude Mythos 5.1
Hacker News · 08-31中立 100%原文 (開新分頁,連回Hacker News貼文)
Breaking Claude Code Opus 5 Auto Mode
evidence refs: evidence/2026-09-29.jsonl#l1, evidence/2026-09-29.jsonl#l2, evidence/2026-09-29.jsonl#l3 …共 20 筆
GPT-6 Sol
openai/gpt-6-solOpenAI
引用來源 21 則
來源分布 Reddit 10 · Hacker News 8 · X 3 ・ 逐則態度由評審團多數決判定
X · 09-29中立 100%原文 (開新分頁,連回X貼文)
Monday: 'I'll keep up with AI news.' Friday: my last brain cell trying to remember which one is Sonnet 5.5, Opus 5.5, GPT-5.5 and GPT-6 Sol 🤯 https://t.co/P3TjvJRv2d
Hacker News · 09-22中立 100%原文 (開新分頁,連回Hacker News貼文)
GPT-6 Sol
Reddit · 09-22正面 50%原文 (開新分頁,連回Reddit貼文)
PSA- GPT-6 Sol Just Launched. ChatGPT Plus Still Can’t Use GPT-6 in Regular Chat. This is very bizarre considering GPT-6 Sol is explicitly cheaper, by 50%, than 5.6 Sol. Why are th…
談到的面向:智能正面速度負面Token 效率負面
Reddit · 09-27中立 75%原文 (開新分頁,連回Reddit貼文)
When Opus 5.5 works on code written by GPT 5.6 Sol (actual screenshot) While your overall statement is not bad, I am just not going to take advise from a person that had to remove …
X · 09-29負面 100%原文 (開新分頁,連回X貼文)
@adhamuxi Yes. And gpt-6-sol is haiku 3 level of intelligence. So I’m over to Claude for the remainder of 2026
談到的面向:智能負面
X · 09-29正面 100%原文 (開新分頁,連回X貼文)
GPT-6 Sol is now available in PowerPoint, helping you tell your story with clearer structure, stronger polish, and brand-ready visuals in all your presentations. Another day, anoth…
談到的面向:智能正面
Reddit · 09-22正面 100%原文 (開新分頁,連回Reddit貼文)
GPT-6 Sol is a disappointment according to AA Intelligence Index! Astra and Fable 5.1 also don't have any use case left after Opus 5.5 The cost cutting is the most significant thin…
談到的面向:智能正面CP 值正面
Hacker News · 09-25正面 100%原文 (開新分頁,連回Hacker News貼文)
DeepSeek beats GPT-6 Sol in autonomous drug development
談到的面向:智能正面
Hacker News · 09-24中立 100%原文 (開新分頁,連回Hacker News貼文)
GPT-6 Sol vs. Luna
Hacker News · 09-22中立 100%原文 (開新分頁,連回Hacker News貼文)
GPT-6 Sol (Max) Intelligence, Performance and Price Analysis
Hacker News · 09-23中立 100%原文 (開新分頁,連回Hacker News貼文)
GPT-6 Sol might acutally be GPT-6 Terra
Hacker News · 09-23中立 100%原文 (開新分頁,連回Hacker News貼文)
"I know I am GPT-6. I cannot confirm whether this session uses GPT-6-Sol."
Reddit · 09-22中立 100%原文 (開新分頁,連回Reddit貼文)
Gpt 6 sol and luna   submitted by   /u/Rollertoaster7   to   r/codex [link]   [comments]
Reddit · 09-23負面 100%原文 (開新分頁,連回Reddit貼文)
Okay, they literally cut our quota by half. GPT 6 Sol got hit too. yeah gonna be honest, I’m not seeing better usage vs 5.6 and it’s marketed as being half the cost. Opus 5.5 is bl…
談到的面向:智能負面Token 用量負面CP 值負面
Reddit · 09-23正面 75%原文 (開新分頁,連回Reddit貼文)
GPT-6 Sol is just GPT-6 Terra renamed as Sol This is the only reason why 5.6 scores better in benchmarks. EDIT: I actually love this approach more than having 4 models better than …
Reddit · 09-22中立 100%原文 (開新分頁,連回Reddit貼文)
Introducing GPT-6 Sol and Luna   submitted by   /u/DemiPixel   to   r/OpenAI [link]   [comments]
Reddit · 09-22負面 100%原文 (開新分頁,連回Reddit貼文)
GPT 6 Sol and Luna Prices! WTF!?   submitted by   /u/nofiler   to   r/codex [link]   [comments]
談到的面向:CP 值負面
Reddit · 09-22正面 75%原文 (開新分頁,連回Reddit貼文)
Introducing GPT-6 Sol and Luna Anyone else remember when it took 6 months for new models to come out? cost cut in half, AND better performance in benchmarks? there is no wall at al…
談到的面向:智能正面CP 值正面
Hacker News · 09-11中立 100%原文 (開新分頁,連回Hacker News貼文)
GPT-6-sol appeared on OpenAI API
Hacker News · 09-22正面 100%原文 (開新分頁,連回Hacker News貼文)
GPT-6 Sol and Luna push the cost-efficiency frontier
談到的面向:CP 值正面
Reddit · 09-07評審團平手原文 (開新分頁,連回Reddit貼文)
Usage tip: “GPT-6 Astra on low performs better than GPT-5.6 Sol on high.” To ease the load on our GPUs* Yeah, and it costs like Sol on max. What about of weekly usage? Is it better…
談到的面向:Token 用量負面
evidence refs: evidence/2026-09-29.jsonl#l109, evidence/2026-09-29.jsonl#l110, evidence/2026-09-29.jsonl#l111 …共 21 筆
Grok 4.7
x-ai/grok-4.7xAI
引用來源 35 則
來源分布 Hacker News 12 · X 12 · Reddit 11 ・ 逐則態度由評審團多數決判定
Hacker News · 09-21中立 100%原文 (開新分頁,連回Hacker News貼文)
Grok 4.7
Reddit · 09-21正面 75%原文 (開新分頁,連回Reddit貼文)
Did they give up on Grok 4.7? update: its out! I’ve been waiting many a days now for this release; I know it’s not the next Fable or Astra but I’m more than ok with a ”better” Grok…
談到的面向:智能正面
X · 09-29中立 100%原文 (開新分頁,連回X貼文)
Grok 4.7 just hit Bedrock. https://t.co/hNYDSE6Lw2
X · 09-29中立 100%原文 (開新分頁,連回X貼文)
Grok 4.7 is now on Amazon Bedrock https://t.co/pKPdMRmDHQ
X · 09-29負面 100%原文 (開新分頁,連回X貼文)
I also used Grok 4.7 today… and I literally had to use sarcasm for it to understand what I wanted. 😂 Absolutely surreal.
談到的面向:智能負面
X · 09-29正面 100%原文 (開新分頁,連回X貼文)
Grok 4.7 Build executing highly efficient and truth based Intelligence asks over thousands of private files. The default to truth is kind of important when running financial analys…
談到的面向:智能正面
X · 09-29正面 75%原文 (開新分頁,連回X貼文)
From my understanding they often have multiple iterations of models going on at the same time, kind of like what Elon does with Grok, 4.6 4.7 4.8 all done in parallel, don't count …
X · 09-29評審團平手原文 (開新分頁,連回X貼文)
@HermesAgentTips Nah. Use Grok 4.7. Much better.
X · 09-29負面 75%原文 (開新分頁,連回X貼文)
@FelipeFr1702 That’s hilarious 😂 When sarcasm becomes part of the prompting strategy, you know the model has developed a very specific personality. Grok 4.7 can be brilliant, but …
談到的面向:智能負面
Hacker News · 09-18中立 100%原文 (開新分頁,連回Hacker News貼文)
Grok Voice Transcribe 2.0
Hacker News · 09-21中立 100%原文 (開新分頁,連回Hacker News貼文)
Grok 4.7 Intelligence, Performance and Price Analysis
Reddit · 09-21負面 100%原文 (開新分頁,連回Reddit貼文)
Grok 4.7 is about 2.5 times as expensive as 4.6 I have been running the exact same task for 2 days now with a Cursor Ultra plan, so I can make a really good comparison between Grok…
談到的面向:CP 值負面
Reddit · 09-21中立 100%原文 (開新分頁,連回Reddit貼文)
Grok 4.6 vs 4.7 Hey i want to share my experiment with the new grok. I wanted to test them a little and quickly and gave them both same prompt, both on extra high. I wanted to make…
Hacker News · 09-28中立 100%原文 (開新分頁,連回Hacker News貼文)
Extrinsic World Modeling with Opus, Astra and Grok
Hacker News · 09-22中立 100%原文 (開新分頁,連回Hacker News貼文)
Errand – open-source Grok Bot and Muse alternative, built in a week
Reddit · 09-23負面 100%原文 (開新分頁,連回Reddit貼文)
Grok 4.7 Performance in Cursor 4.7 hasnt been noticeable technical improvement in my use & it seems to eat usage faster. no longer using it for cursor model type tasks. it's suppos…
談到的面向:智能負面速度負面Token 用量負面CP 值負面
X · 09-29中立 75%原文 (開新分頁,連回X貼文)
Grok 4.7 hits #1 on Artificial Analysis Cyber Index
Reddit · 09-22負面 100%原文 (開新分頁,連回Reddit貼文)
Grok 4.7 is officially more c#nsored than Claude & ChatGPT! PROOF: https://preview.redd.it/2wcfcffle3rh1.png?width=760&format=png&auto=webp&s=3eb778e8796092a3709d45d79f…
Reddit · 09-21中立 100%原文 (開新分頁,連回Reddit貼文)
Grok 4.7 is out   submitted by   /u/Existing_Hat_1064   to   r/grok [link]   [comments]
Reddit · 09-21負面 100%原文 (開新分頁,連回Reddit貼文)
I just tried Grok 4.7 and it is... When i find a local llm that will take images and write spicy llm prompts im probably done with grok. Grok was useful when it prouded itself on b…
Reddit · 09-21負面 100%原文 (開新分頁,連回Reddit貼文)
People think we will have big AI Releases this week. Grok 4.7 was released minutes ago If they don't fucking fix opus and sonnet I swear to.... How’s pacing the frontier going, lad…
談到的面向:智能負面
X · 09-29負面 100%原文 (開新分頁,連回X貼文)
Grok 4.7 has severe amnesia & is mispronouncing words like super bad
談到的面向:智能負面
Reddit · 09-21負面 75%原文 (開新分頁,連回Reddit貼文)
Introducing Grok 4.7 Twice as fast at half the price is a pretty aggressive claim I'm more curious how Grok 4.7 actually performs inside Cursor . I still want Composor 3 SpaceXAI c…
X · 09-29負面 100%原文 (開新分頁,連回X貼文)
What was the point of releasing grok 4.7, knowing damn well it’s terrible?
談到的面向:智能負面
Reddit · 09-21正面 100%原文 (開新分頁,連回Reddit貼文)
Grok 4.7 benchmarks Why is Sol on here and not Astra That's decently impressive for the price. trying to do rough price matching
談到的面向:智能正面CP 值正面
Reddit · 09-21負面 100%原文 (開新分頁,連回Reddit貼文)
Introducing Grok 4.7 Can I goon with it, Elon? If not, fuck off. for coding and knowledge meaning its useless for rp and writing, sigh... wait for 4.8 then... It’s good for writing…
談到的面向:智能負面
X · 09-29中立 75%原文 (開新分頁,連回X貼文)
Grok 4.7 xHigh is #1 on the Artificial Analysis Cyber Index. It reportedly beats Fable 5.1 Max, Opus 5.5, Astra 6, GPT-6 and other frontier models. AI coding is getting seriously c…
X · 09-29評審團平手原文 (開新分頁,連回X貼文)
Grok 4.7 xHigh가 Artificial Analysis Cyber Index에서 1위를 차지했습니다. 기업용 사이버 방어 분야에서 Fable 5.1 Max, Opus 5.5, Astra 6, GPT-6 등을 모두 앞선 성능을 보여줬어요.
談到的面向:智能正面
Hacker News · 09-03中立 75%原文 (開新分頁,連回Hacker News貼文)
Grok outage
Hacker News · 09-18中立 100%原文 (開新分頁,連回Hacker News貼文)
Uncle Bob Martin's UML Tool to Manage Grok AI Agents
Hacker News · 09-15中立 100%原文 (開新分頁,連回Hacker News貼文)
Grok Bot Galaxy
Hacker News · 09-17中立 100%原文 (開新分頁,連回Hacker News貼文)
Show HN: Jev routing coding tasks to Grok Build or Codex Astra
Hacker News · 09-03中立 100%原文 (開新分頁,連回Hacker News貼文)
ChatGPT, Claude, and Grok Are Down
Hacker News · 09-13中立 100%原文 (開新分頁,連回Hacker News貼文)
Expanding model choice in Copilot with Grok
Hacker News · 09-03中立 100%原文 (開新分頁,連回Hacker News貼文)
Ask HN: Why were OpenAI, Claude, and Grok simultaneously down?
evidence refs: evidence/2026-09-29.jsonl#l153, evidence/2026-09-29.jsonl#l154, evidence/2026-09-29.jsonl#l155 …共 35 筆
分數 0~100,總分是各面向依權重加權後重歸一。標示資料不足 代表該面向沒有被討論過,不等於 50 分。本次評分由 LLM 評審團 ×4 以多數決判定。
怎麼看這張表
- 總分
- 各面向依權重加權後重歸一(品質類權重最高,價格與速度次之), 所以總分含蓋社群在乎的面向, 不等於單一能力值。
- 維度切換
- 欄位由資料的
meta.dimensions驅動,切換後依該面向重新排序; 沒有討論過的面向一律排在最後,不當成 50 分。 - 樣本數與信心
N是該模型被蒐集到的貼文筆數; 樣本信任度是正面率的 Wilson 95% 下界。 N 低於 20 標「僅供參考」, 單一面向討論數低於 10 則標「樣本不足」。- 引用來源
- 每列底下的「引用來源」展開後是逐則貼文(來源、時間、評審團態度、原文連結), 可直接連回原文查核。
資料產出於 2026-09-30 13:19(台北,29 小時前)·評審團 ×4 · 共 8 個模型入榜
想知道這些數字背後的方法與限制 →讀這個章節的技術文章