Wink Pings

Stanford cracked resampling with two models auditing each other? Those two are not in the paper at all

A tweet claims that GPT-6 Astra and Claude Opus 5.5 broke through the resampling wall, with a 28-point improvement on MATH500. But people who read the original paper found that the experiment actually used small open-source models of 1.5B parameters. The method might work, but the hype got ahead of the facts.

A tweet has gone viral: A Stanford research team found a way to make GPT-6 Astra and Claude Opus 5.5 work together and directly broke through the resampling wall. The results are impressive: accuracy on MATH500 jumped from 29.6% to 58.0%, and on AMC-2023 it increased by 7.5 percentage points, all without using any reward model.

What is the resampling wall? Most developers use the test-time compute approach, which means running the same model 32 times blindly, burning through a huge number of tokens, only to end up falling into the same 2-3 error traps over and over again. The core idea of Stanford's paper (arXiv:2608.05643) is to split the computation between two different models, each handling its own task.

According to the tweet, the process has four steps:

- Breadth-first search: GPT-6 Astra uses high-entropy exploration (T=0.8) to find different problem-solving paths

- Step-by-step auditing: Claude Opus 5.5 checks each step line by line at T=0.0 to locate logical errors

- Local repair: Opus only rewrites the broken step instead of restarting the entire reasoning chain from scratch

- Verifier-free consensus: The result is determined by majority voting, no reliance on PRMs which are prone to drift

T here is the temperature parameter. T=0.8 means more randomness and diversity; T=0.0 means completely deterministic. PRM stands for process reward model, which scores each reasoning step, and this process bypasses it entirely.

The image in the video shows the complete workflow and metrics:

Per the tweet, the metrics for this workflow are: 91.2% recovery rate for detected logical flaws, and a regression rate on correct branches of less than 1.8%. Astra explores paths, Opus fixes them. Instead of wasting tokens on blind attempts, you now spend resources on verified reasoning depth.

But there is a crack in this story.

A netizen named Kagi commented in the thread: After reading the paper, there is no mention of Opus 5.5 or GPT-6 Astra. The experiments were run on open-source models ranging from 1.5B to 8B parameters, such as Qwen2.5 and LLaMA-3.1. The 58% accuracy on MATH500 was achieved by a 1.5B model.

If this is true, the tweet just slapped the names of two popular models onto the paper's results. Actually, this kind of packaging is unnecessary: getting nearly double the accuracy on a 1.5B small model through this framework is a more counter-intuitive and significant finding than two large models working together. If even small models can benefit from this method, it proves that the method itself is universally applicable.

The original tweeter marfin also shared a long article explaining the complete architecture of Test-Time Compute Engineering, covering engineering details such as search tree topology, budget forcing, and early stopping. Those who want to implement this method can read it here: [Test-Time Compute Engineering](http://x.com/i/article/2084413967930200064). But note that this long article also does not mention the specific models used in the paper.

What is worth remembering from this incident is not the data, but how it spread: an effective method was packaged into hype about a combination of popular models. In the AI community, don't retweet impressive-looking data right away — go check the original paper. Even if the original result came from small models, it might still be the real treasure.

发布时间: 2026-10-05 08:00