Skip to content
Archive

Post

Back to deliverables

X Xudong07452910
Xudong Han
@Xudong07452910

过去两个月,我们一直在闷声干一件事:教 AI 预测未来。 今天可以摊牌了 🎉 在 FutureX,一个专门考察"预测还没发生的事"的实时榜单上,我们的长程预测 agent 拿下: 🥇 第 1 名(基座 Kimi K3) @Kimi_Moonshot 🥉 第 3 名(基座 DeepSeek V4 Pro)@deepseek_ai 7️⃣ 第 7 名(基座 MiniMax M3)@MiniMax_AI 同一套框架,换三个模型当"大脑",全部进前十。 感谢团队成员 @Bruce_Lee1207 一起努力的付出 。 先回答一个你们一定会问的问题:现在榜单满天飞,FutureX 凭什么值得认真对待? 1️⃣ 出品方是 ByteDance Seed 联合 Stanford、Princeton、复旦,学界业界联合出品,论文在 arXiv 公开; 2️⃣ 它可能是唯一一个"打榜"在逻辑上不成立的榜单:题目全是尚未发生的真实事件,agent 先交预测,官方等事件落地后按真实结果判分。答案在出题时根本不存在,所以没有题库可背,没有答案分布可过拟合,想提分只有一条路:真的更会预测; 3️⃣ 同台的是谁:OpenAI、Google、xAI 的 Deep Research 系 agent 都被系统评测过,H2O.ai 登顶时还专门发了官方博客。 再说我们的 agent 在做什么。你可以把它想象成一个不知疲倦的分析师: 🔍 自己上网检索,多信源交叉验证 🧠 把证据沉淀成"信念",有新信息就更新判断 🎲 最后给的不是一句"可能吧",而是一个校准过的概率 ⚖️ 多个子分析师意见相左时,还有裁决机制出面对齐 ...... 为什么执着于预测?因为预测未来是智能最诚实的试金石,它逼着模型把检索、推理、常识、对不确定性的拿捏全部亮出来,一点都藏不住。 这只是开始。方法细节,后续成绩和最后的产品,会陆续放出来。

photo

· 25K Views

16 Reposts 4 Quotes 112 Likes 82 Bookmarks
replies reposts likes
18 replies collected
李韭二 @li9292 · 13K

@Xudong07452910 @Kimi_Moonshot @deepseek_ai Looking forward to your product!

original · zh

@Xudong07452910 @Kimi_Moonshot @deepseek_ai 期待老师的产品!

1 4
Xudong Han @Xudong07452910 ·

@Kimi_Moonshot @deepseek_ai DeepSeek V4 Pro is incredibly powerful. I'm really looking forward to its official release; I feel like its capabilities will take another big leap forward.

original · zh

@Kimi_Moonshot @deepseek_ai DeepSeek V4 Pro太抗打了,很期待它的正式版发布,感觉能力又得上一个大台阶。

4
F @FrankFred834567 ·

@Xudong07452910 @Kimi_Moonshot @deepseek_ai This ranking list is quite prestigious because the answers don't exist when the questions are created, so it's impossible to overfit by simply filling in a question bank.

original · zh

@Xudong07452910 @Kimi_Moonshot @deepseek_ai 这个榜单含金量够硬,因为答案在题出的时候还不存在,没法靠刷题库过拟合。

1 3
Fred-New @Fred834567 ·

@Xudong07452910 @Kimi_Moonshot @deepseek_ai I'd like to know how the adjudication mechanism aligns the opinions of multiple analysts. Is it weighted voting or tiered arbitration?

original · zh

@Xudong07452910 @Kimi_Moonshot @deepseek_ai 想知道裁决机制怎么对齐多分析师意见的,是加权投票还是分层仲裁?

1 1
momoai沫沫🫧 @momoai_daily ·

@Xudong07452910 @Kimi_Moonshot @deepseek_ai Kimi is so talented

original · zh

@Xudong07452910 @Kimi_Moonshot @deepseek_ai kimi这么有实力

1 1
Jason huang @bbs2i58 ·

@Xudong07452910 @Kimi_Moonshot @deepseek_ai Predicting things that haven't happened yet is a false proposition; it's similar to fortune-telling. I did a similar investment prediction, also using an LLM model, but the model comparison results aren't out yet.

original · zh

@Xudong07452910 @Kimi_Moonshot @deepseek_ai 预测还没发生的事其实是个伪命题,这个和算命差不多,我做过一个类似的投资预测,也是接LLM,模型对比结果还没出来

1 1
Cunxiang Wang @CunxiangWang ·

@Xudong07452910 @Kimi_Moonshot @deepseek_ai Is there any way to prove that "prediction" is closely related to model ability, rather than relying more on luck? For example, the results of GLM-5.2 with different harnesses can differ by 26 points, which feels more like random guessing than ability. Also, GLM-5.1 is significantly better than GLM-5.2 on the leaderboard, which doesn't quite meet public expectations. I'm also thinking about issues related to academic prediction.

original · zh

@Xudong07452910 @Kimi_Moonshot @deepseek_ai 有没有办法证明“预测”这个事是和模型能力紧密相关,而不是更靠运气呢?比如GLM-5.2配不同harness的结果能相差26分,感觉更像是一种random guess而不是能力。而且榜单里glm-5.1明显比glm-5.2好得多,这也不太符合大众预期。我也在思考agentic prediction相关的问题

4 5
yishan @tspy ·

@Xudong07452910 @Kimi_Moonshot @deepseek_ai But current LLMS are all based on existing data and make weighted judgments, making it difficult to make truly innovative predictions.

original · zh

@Xudong07452910 @Kimi_Moonshot @deepseek_ai 但现在的llms都是基于已有数据做的一个权重的判断,很难做出真正创新性的预判

2 1
Oldeng @dengdry ·

@Xudong07452910 @Kimi_Moonshot @deepseek_ai Holy crap, quietly working on a huge project!

original · zh

@Xudong07452910 @Kimi_Moonshot @deepseek_ai 卧槽,闷声搞大项目啊

1 1
Tz @Tz_2022 · 25K

@Xudong07452910 @Kimi_Moonshot @deepseek_ai This product can now be officially put on the agenda... x.com/Tz_2022/status/208285574…

original · zh

@Xudong07452910 @Kimi_Moonshot @deepseek_ai 可以把这个产品正式提上日程了。。。 https://x.com/Tz_2022/status/2082855746505503178/video/1?s=46

2 2
诺鸭船长3 @noahduck283 · 19K

@Xudong07452910 @Kimi_Moonshot @deepseek_ai Wow, I really have to pay to experience Kimi.

original · zh

@Xudong07452910 @Kimi_Moonshot @deepseek_ai 牛的,真得付费体验一下kimi了

1 1
Wukong @WukongNumber1 ·

@Xudong07452910 @Kimi_Moonshot @deepseek_ai It was ready before the World Cup!

original · zh

@Xudong07452910 @Kimi_Moonshot @deepseek_ai 世界杯之前出好了!

1 1
Wang chen_AI26 @wangchen_AI26 ·

@Xudong07452910 @Kimi_Moonshot @deepseek_ai "Predicting the future is the most honest touchstone for intelligence." Prediction is not like manipulating rankings; there are no standard answers to overfit. In the end, only genuine ability matters. Furthermore, the fact that the same framework, with three different bases, all made it into the top ten indicates that the ceiling lies not in the model but in the architectural design.

original · zh

@Xudong07452910 @Kimi_Moonshot @deepseek_ai "预测未来是智能最诚实的试金石",预测不像刷榜,没有标准答案可以过拟合,最后只能靠真本事。 而且同一套框架换三个底座都进前十,说明天花板不在模型而在架构设计。

1
Kelvin @ai_Goge ·

@Xudong07452910 @Kimi_Moonshot @deepseek_ai Save this and study it carefully.

original · zh

@Xudong07452910 @Kimi_Moonshot @deepseek_ai 这个的收藏起来好好研究

1 1
Lam @Lamudpz ·

@Xudong07452910 @Kimi_Moonshot @deepseek_ai Can you predict the polymarket?

original · zh

@Xudong07452910 @Kimi_Moonshot @deepseek_ai 能预测polymarket吗(

1
指月图鉴 @iChineseWhisper ·

@Xudong07452910 @Kimi_Moonshot @deepseek_ai Prediction is intelligence.

original · zh

@Xudong07452910 @Kimi_Moonshot @deepseek_ai 预测即智能。

1
Erwin @ErwinWu000 ·

@Xudong07452910 @Kimi_Moonshot @deepseek_ai Brother Han is making a fortune in silence!!! We're eagerly awaiting your agent.

original · zh

@Xudong07452910 @Kimi_Moonshot @deepseek_ai 韩哥闷声发大财!!!翘首以盼你的Agent了

1 1

8 replies whose parent comment X withheld

Cunxiang Wang @CunxiangWang ·

@Xudong07452910 @Kimi_Moonshot @deepseek_ai Yes, it's definitely a combination of harness + LLM. However, I think that the stable and reproducible differences in a fixed harness might be random guessing. For example, in a dice game, agent A likes to bet heavily on 3, while agent B likes to bet on 4. If there are a few rounds with more 3s than 4s, then agent A will appear to score much higher than agent B, no matter how many times you test it again. However, real-world scenarios are definitely much more complex.

original · zh

@Xudong07452910 @Kimi_Moonshot @deepseek_ai 嗯,肯定是harness+LLM的组合。不过我觉得固定harness差异稳定可复现,也可能是random guess,打个比方,掷骰子的游戏,agentA就喜欢猛压3,agentB就喜欢压4,而正好几局里3多4少,那agentA就会显得比agentB分数高很多,无论你让它再回头测几次都这样。不过现实场景肯定复杂得多

1 1
Cunxiang Wang @CunxiangWang ·

@Xudong07452910 @Kimi_Moonshot @deepseek_ai I think we might need to increase the number of events by 10 times to eliminate randomness. If there is still a significant difference in scores at that point, I think it would be very convincing.

original · zh

@Xudong07452910 @Kimi_Moonshot @deepseek_ai 我觉得可能还是得把event数再变大10倍,消除随机性,如果那时候还很有分差,我觉得就很有说服力了。

2 1
yishan @tspy ·

@Xudong07452910 @Kimi_Moonshot @deepseek_ai I understand. Prediction is about judging the subsequent development trend and possible changes of existing events.

original · zh

@Xudong07452910 @Kimi_Moonshot @deepseek_ai 明白了,预测就是判断既有事件的后续发展趋势及可能发生的变化

1
Xudong Han @Xudong07452910 ·

@CunxiangWang @Kimi_Moonshot @deepseek_ai Yes, the number of problems on this benchmark isn't very large right now, but there aren't any better alternatives. It seems like maintaining this leaderboard is very costly; the problems need to be updated every few days, which might be why they can't increase the number of problems and focus on selecting the best ones.

original · zh

@CunxiangWang @Kimi_Moonshot @deepseek_ai 是嘞,目前这个benchmark的题目数量不是很多,但也没有其他更好的选择。感觉这个榜单维护成本很高,每隔几天就要更新题目,可能这也是导致他们没法增加题目数量,尽量精选题目。

1
Xudong Han @Xudong07452910 ·

@CunxiangWang @Kimi_Moonshot @deepseek_ai That makes sense. We still need to do many more experiments to observe these phenomena. Thank you for your observations and questions.

original · zh

@CunxiangWang @Kimi_Moonshot @deepseek_ai 有道理的,我们还得继续做很多实验是观察这些现象,谢谢您的观察和疑问。

1 2
Cunxiang Wang @CunxiangWang ·

@Xudong07452910 @Kimi_Moonshot @deepseek_ai No, I have to thank you guys for the bench, but I've been thinking about these issues myself lately.

original · zh

@Xudong07452910 @Kimi_Moonshot @deepseek_ai 没有,得感谢你们的bench,只是我最近也在思考这些问题。

1 1
Xudong Han @Xudong07452910 ·

@CunxiangWang @Kimi_Moonshot @deepseek_ai You're welcome. I look forward to exchanging ideas more often in the future.

original · zh

@CunxiangWang @Kimi_Moonshot @deepseek_ai 客气啦,期待以后多多交流。

1
Lam @Lamudpz ·
1

These were collected in full; the comment they answer was not returned by X.