Benchmark
Graphwalks BFS <128k
reasoning
spatial_reasoning
text
A graph reasoning benchmark that evaluates language models' ability to perform breadth-first search (BFS) operations on graphs with context length under 128k tokens, returning nodes reachable at specified depths.
语言EN
满分1
参评模型7