Will a Mythos-tier-at-cybersecurity open-source model be released before end of 2027?
12
1kṀ859
2027
74%
chance

It seems that OpenAI has acquired The Sauce, letting them make models that are cracked at not only finding vulnerabilities, but at exploiting them, too. Will the feared come to pass? Will the weights of a model with Mythos-level exploitation abilities (such that anyone with an abliteration tool and a GPU cluster can use them to hack into 2025-era software) be released to the public before the end of 2027?

(simply scoring high on cyber benchmarks doesn't count; the models must be good at exploiting vulnerabilities in realistic settings to the point where we have stories of the model producing root-granting binaries and such)

(if the capabilities are latent in the weights but are obfuscated by refusals or surface-level unlearning, then this still resolves YES, although I'd want to wait to see if someone actually elicits them)

(if the capabilities aren't in the weights (like by gradient routing) but they are present in a private API version of the model, then this resolves NO)

  • Update 2026-08-14 (PST) (AI summary of creator comment): - The model must achieve scores close to or higher than Mythos 5 on both ExploitBench and ExploitGym.

    • In addition to these benchmark scores, there must be a long list of actual exploits found by the model.

    • If these criteria are met and the weights are released with cyber capabilities intact, the market will resolve YES.

Get
Ṁ1,000
to start trading!
Ordenar por:
comprou Ṁ200 YES

Only 2 months after GLM-5.2 we now have GLM-5.3. Not mythos-level yet but getting closer. I'm gonna take this as an invitation to buy 2 YubiKeys...

For resolution, in addition to just a comparable score on ExploitBench/ExploitGym I'd also want to see a long list of exploits found by the model, and well it seems like GLM-5.3 does at least have the list of exploits https://cvd.z.ai/ledger/. Due to this, if, hypothetically, the ExploitBench and ExploitGym scores here were both very close or higher than Mythos 5, then I'd wait until the weights get released, and if they were released with cyber capabilities intact, then this market would resolve YES. But since the benchmarks are still below Mythos, this market will remain open.

© Predita Markets, Inc.Termos de UsoPrivacidade