# Claude Mythos long-context abduction > The long-context score is the most relevant, because the survey identified long, dispersed-evidence narrative abduction (True Detective, MuSR) as the main failure mode of current models. Mythos is reported at 80.0% on GraphWalks BFS, described as nearly four times GPT-5.4's score on million-token reasoning. If that long-context strength is real and transfers, it targets one of the two things that break abduction in the survey's analysis. But GraphWalks is graph traversal, not abductive hypothesis formation, so this is suggestive at best — the capacity to hold and traverse a long context is necessary for long-narrative abduction but not sufficient for it, since the hard part is inferring the hidden explanation, not tracking the evidence.