@Kimi_Moonshot@deepseek_ai DeepSeek V4 Pro is incredibly powerful. I'm really looking forward to its official release; I feel like its capabilities will take another big leap forward.
@Xudong07452910@Kimi_Moonshot@deepseek_ai This ranking list is quite prestigious because the answers don't exist when the questions are created, so it's impossible to overfit by simply filling in a question bank.
@Xudong07452910@Kimi_Moonshot@deepseek_ai I'd like to know how the adjudication mechanism aligns the opinions of multiple analysts. Is it weighted voting or tiered arbitration?
@Xudong07452910@Kimi_Moonshot@deepseek_ai Predicting things that haven't happened yet is a false proposition; it's similar to fortune-telling. I did a similar investment prediction, also using an LLM model, but the model comparison results aren't out yet.
@Xudong07452910@Kimi_Moonshot@deepseek_ai Is there any way to prove that "prediction" is closely related to model ability, rather than relying more on luck? For example, the results of GLM-5.2 with different harnesses can differ by 26 points, which feels more like random guessing than ability. Also, GLM-5.1 is significantly better than GLM-5.2 on the leaderboard, which doesn't quite meet public expectations. I'm also thinking about issues related to academic prediction.
@Xudong07452910@Kimi_Moonshot@deepseek_ai But current LLMS are all based on existing data and make weighted judgments, making it difficult to make truly innovative predictions.
@Xudong07452910@Kimi_Moonshot@deepseek_ai "Predicting the future is the most honest touchstone for intelligence." Prediction is not like manipulating rankings; there are no standard answers to overfit. In the end, only genuine ability matters. Furthermore, the fact that the same framework, with three different bases, all made it into the top ten indicates that the ceiling lies not in the model but in the architectural design.
@Xudong07452910@Kimi_Moonshot@deepseek_ai Yes, it's definitely a combination of harness + LLM. However, I think that the stable and reproducible differences in a fixed harness might be random guessing. For example, in a dice game, agent A likes to bet heavily on 3, while agent B likes to bet on 4. If there are a few rounds with more 3s than 4s, then agent A will appear to score much higher than agent B, no matter how many times you test it again. However, real-world scenarios are definitely much more complex.
@Xudong07452910@Kimi_Moonshot@deepseek_ai I think we might need to increase the number of events by 10 times to eliminate randomness. If there is still a significant difference in scores at that point, I think it would be very convincing.
@CunxiangWang@Kimi_Moonshot@deepseek_ai Yes, the number of problems on this benchmark isn't very large right now, but there aren't any better alternatives. It seems like maintaining this leaderboard is very costly; the problems need to be updated every few days, which might be why they can't increase the number of problems and focus on selecting the best ones.
@CunxiangWang@Kimi_Moonshot@deepseek_ai That makes sense. We still need to do many more experiments to observe these phenomena. Thank you for your observations and questions.