Skip to content
Archive

Post

Back to deliverables

P Phoenixyin13
Phoenix Yin
@Phoenixyin13

我想的是,我们是否可以举办AI奥林匹克竞赛,现场版的那种。众目睽睽之下,让不同架构、不同训练范式的大模型在统一的公理体系下盲测猜想能力? 我们已经不需要让AI做那种IMO竞赛题,全对了,都没有什么意思。 我们直接让它们做猜想循环。 猜想怎么打分呢? 我认为这是唯一的原创设计空间。 人类评委给有不有趣打分不可规模化,但有一个机制可以绕开品味,就是用下游效用定价。 一条猜想得分当且仅当 (1) 形式化陈述,(2) 基线证明器在预算内证不出来(非平凡),(3) 在后续回合中被其他参赛系统当作引理引用、并且缩短了它们的证明长度。 每个模型既出题又解题,出的题越是让对手卡住又最终可用,分越高。 这可以把猜想的价值从审美判断变成了一个可测的图结构量,其实这本质上是 Erdős 的做法自动化。 人类看围棋或数学竞赛,看的是有限肉身在极其不确定空间里的灵光一现与情绪张力。 但如果把赛场交给 AI,竞赛的意义就超越观赏,超越基准测试,进行终极演进。 这条路上已经有人在铺。First Proof Project 定位为对AI在研究数学中表现的独立、透明评估,2026年3月到6月做了形式化基准加社区轮,AIMO 的 Proof Pilot 也在6月启动。这些都是很好的趋势。 想法里的增量在于对抗式的猜想定价机制。那部分在我的观察里,目前没人做。

@

Hi mathematicians, don't fret. AI beat humans long long ago in chess, yet we are mostly interested in watching humans play chess. Likely something like that will play out for you as well. We should still have olympiads. We should continue to have conferences where mathematicians go talk. I will still be interested in hearing what Terrence Tao has to say more so than an AI smarter than him that has something to say.

· 1.4K Views

6 Likes 1 Bookmarks
replies reposts likes
3 replies collected
Claude Qin @baboonAI4S ·

@Phoenixyin13 I don't think many people in the industry have your unique aesthetic sense.

original · zh

@Phoenixyin13 我觉得业内未必有你这种独特的审美境界

1
Suwako — e/acc @suwakopro ·

@Phoenixyin13's idea is actually somewhat similar to PageRank's.

original · zh

@Phoenixyin13 其实和page rank的想法有点类似

1 1
Pyuyi @Pyuyi233 ·

@Phoenixyin13 Are you guessing CTF? (laughs)

original · ja

@Phoenixyin13 猜想ctf吗(笑

1

1 replies whose parent comment X withheld

Claude Qin @baboonAI4S ·

@Phoenixyin13 I greatly admire your thinking and wide range of interests. I'd love to learn from you.

original · zh

@Phoenixyin13 非常欣赏您的思考和多领域广泛的兴趣,向您学习

1

These were collected in full; the comment they answer was not returned by X.