Will I solve an unsolved math problem with AI in August?
209
1kṀ47k
Aug 31
3%
chance

Should be a top 20 math problem of its math subfield, so nothing crazy but a couple people care a fair bit about it.

I am just starting a project to see how doable this is at scale, may drop it if it doesn’t go anywhere. Tips welcome.

By ‘solve’ I mean AI is very confident that the solution is correct. A mathematician needn’t sign off on it. If you think my AI ends up wrong about a solution I’d be happy to bet on it.

  • Update 2026-08-12 (PST) (AI summary of creator comment): The creator has specified the exact methodology for defining a math subfield and its top 20 problems:

    • Subfield granularity: Defined at the community level (roughly 300–500 total subfields, e.g., 'spectral geometry' is valid, but 'geometry' or 'zero-sum theory' are not).

    • Subfield selection: Identified by prompting 12 '5.6 sol' and 12 'opus 5' AI subagents. A subfield qualifies if it receives 5+ mentions from either group.

    • Problem ordering: For each subfield, agents will generate the top ~30 problems and vote on their ordering to establish the top problems.

  • Update 2026-08-12 (PST) (AI summary of creator comment): The resolution criteria have been updated regarding what constitutes a 'solved' problem:

    • The proof must be correct to count for a YES resolution.

    • Correctness will be evaluated by a fleet of adversarial verifier subagents (5.6 sol / 5.6 sol ultra).

    • If the proof is ultimately proven incorrect, it will not count.

Get
Ṁ1,000
to start trading!
Ordenar por:
comprou Ṁ10 YES

Hopefully Astra is released before the end of the month 🤣🤣

vendeu Ṁ0 YES

@GazDownright No, definitely not! For one, this wasn't solved with AI but rather computational algorithms. He's also a Regeneron STS finalist so he's pretty much a prodigy anyway.

This isn't an example of AI being cracked, but rather a 17-year-old being cracked.

You should solve the square free Fermat Number question.

For prompting, in my experience when generating new math research it has been very important to make the AI itself believe it can do the task, and to use authoritative language "do this," "work until you've achieved that" to avoid hedging semantics. It might even be better to use a fresh context window and present the problem from first principles--"prove there are no natural a,b,c,n>2 s.t. a^n+b^n=c^n"--instead of presenting a problem as e.g. "the well known Fermat conjecture" because otherwise the AI thinks it can't do it. Obviously, it'll recognize the problem quickly, but it does seem to help regardless.

Solve this; it is of the type I could imagine AI solving: Gilbreath's conjecture - Wikipedia

@JussiVilleHeiskanen seems way too hard to me!

Solve the hailstone conjecture.

preenchido a Ṁ5 YES at 28% order

Can you provide a bit more background in terms of your formal/informal mathematical education?

aberto a Ṁ2,000 YES at 10% order

@dorothydomer 1 semester away from finishing my math&cs undergrad degree, not intending to have to follow any of the math here gpt 5.6 sol will do all the heavy lifting. All the lifting really

an independent fable instance suggested:

On subfield granularity: I think the right level is the community level — X is a subfield if a working mathematician would plausibly write "I work in X" on their homepage, and X has its own survey literature, recurring workshops, and something like 50–500 active researchers. In MSC terms that's roughly the middle tier, giving maybe 300–500 subfields of math total. So: "geometry" is a field, not a subfield. "Spectral geometry" — yes, subfield. "Game theory" is borderline-field; "algorithmic game theory" is a clean subfield. "Zero-sum theory" and "zero-error information theory" are too narrow — those are topics within subfields (additive number theory and information theory respectively). The failure mode you want to guard against is gerrymandering: any solved problem is trivially top-20 of a sufficiently narrow topic, so the subfield definition has to be fixed independently of the problem. The test I'd use: would 10 AIs, asked cold to "list subfields of mathematics" at survey-article granularity, generate this category without being prompted by the problem? Zero-error information theory fails that; spectral geometry passes.

and then i had 12 5.6 sol subagents and 12 opus 5 subagents generate a list from that prompt, and for each of those two groups, subfields with 5+ of the 12 mentioning it counts in the list of subfields, and then i'll have for each subfield a couple agents try to generate the top 30ish problems in it, and then have them all vote on how to order them, to get a fairish agreement on the ordering. they risk to miss some problems, which is a pretty serious problem with this methodology, so lmk if you have another idea

there are like ~430 subfields identified, so that will make like 8600 problems? which is a lot of problems, maybe makes this too easy? thoughts welcome

@Bayesian what's your budget?

@Bayesian I doubt you'll get 8600 problems out of that. Some of those subfields will be quite niche and probably won't have 20 unsolved problems that are recognizable to those in the field. Some of those problems also might be equivalent as well. Where the questions posed by the problems are different but both are only unsolved because of the same missing proof.

@Bayesian maybe this is dumb and certainly it will be uneven, but there is a tree of categories on wikipedia that is sort of not totally arbitrary and or for the nonce

@AviEisenberg a chatgpt pro 20x subscription

@Zeolite i agree this is trouble. i'm aiming for something like min(20, # of meaningful unsolved problems in the subfield)

@JussiVilleHeiskanen there's like 150 name overlaps between my list and wikipedia, claude tells me wikipedia has a hundredish pedagogical/umbrella/historical categories that aren't active areas of research or stuff like "survey methodology" which doesn't seem like it would ahve the kinds of problems i'm looking for

"By ‘solve’ I mean AI is very confident that the solution is correct." Very easy to do. Counterintuitively I suggest using a weaker model.

@ProjectVictory i’m pretty sure this is wrong, 5.6 sol is extremely good at finding flaws in incorrect proofs, older models are much worse

@Bayesian this is right. Since the point of the market is to find the proof that a model would believe to be correct, not a necessarily correct proof, it would be much easier to do with a weaker model.

@ProjectVictory oh i see. Ok ill reframe as, the resolution criterion is tjat the proof has to be correct. My default way of figuring that out is asking a fleet of adversarial verifier 5.6 sol / 5.6 sol ultra subagents. By default i assume Their collective stamp of approval is sufficient, but if the proof ends up being wrong it doesnt count

I only looked through a small sample of those in the group theory section of problem garden, and the commentary on a surprising number of them suggested they were vulnerable. Not sure if neglect keeps fertile or are they being encoraging deceptively to entice eyeballs.

@JussiVilleHeiskanen idk what this means

Not thrilled about that resolution criteria. "Bullshit engine confirms its solution is 100% legit, trust me bro."

@BarryJones me too, do you have a better idea?

@Bayesian maybe just spell out the criteria in more detail. Which model or models at which thinking level have to agree that it's a solution? Or maybe have prop bets or something for how well verified it'll be? Will a math PhD peer review and affirm the solution, etc.

@Bayesian Lean certificate?

Not all top problems have Lean problem statements though. "Lean certificate if there's already an agreed Lean problem statement" might be a good start.

@BarryJones Latest Fable or Sol (or Astra?) release that didn't work on the solution in question needs to review it and say it passes, maybe? With use of sub agents?

@EvanDaniel im trying to seek the low hanging fruits, that’s part of the point of the project, so aiming for hard to formalize but Important problems. So looking at like a bank of lean formalized problems is what i want to avoid

@Bayesian Yeah, I agree that "only include Lean formalized" would change the spirit of the question in an unhelpful way. Hence my suggestion that if there's a formalized problem statement, demand a formalized solution. I definitely agree that unformalized problems is a rich domain!

That said, I suspect a lot of the low hanging fruit here is also low hanging with respect to formalization work; if you find a promising solution, I would definitely recommend seeing if the model can formalize both problem and solution in Lean. I am certainly not arguing to rule things out if the answer is no, though!

@EvanDaniel yeah ok makes sense

© Predita Markets, Inc.Termos de UsoPrivacidade