Markovian ODE-guided scoring can assess the quality of offline reasoning traces in language models

Arghodeep Nandi · Ojasva Saxena · Tanmoy Chakraborty

Video

Paper PDF

Thumbnail of paper pages

Abstract

Reasoning traces produced by generative language models are increasingly used for tasks ranging from mathematical problem solving to automated fact checking. However, existing reasoning evaluation metrics are typically validated using either synthetic perturbations or benchmark-specific correlations. Moreover, the design choices underlying many metrics are often closely tied to the perturbation sets used for their evaluation, raising concerns about their generalizability. Additionally, despite employing multiple perturbation types, prior work frequently relies on binary grading schemes that provide only a coarse assessment of reasoning quality. Consequently, such evaluations offer limited evidence that a metric remains sensitive to realistic degradations in reasoning quality or that it generalizes across reasoning domains. As a result, it remains unclear whether existing metrics capture properties that align with human judgments of reasoning quality. To this end, we introduce MarODE, an offline evaluation framework that assigns quality scores to reasoning traces. Its effectiveness is assessed using metric-agnostic graded perturbations and human judgments, which jointly evaluate the fundamental dimensions of an evaluation metric – goodness and soundness. The approach is grounded in a Markovian formulation of reasoning progression and an ordinary differential equation based characterization of trace dynamics, enabling efficient evaluation of reasoning quality. In a large-scale evaluation, MarODE outperforms existing baselines consistently under Somers’ D correlation. Our results emphasize the value of theory-driven evaluation frameworks as reasoning traces become central to language model-based systems.