From Causal Plausibility to Causal Reliability: Evaluating LLMs as Calibrated Direct Causal-Edge Classifiers
A systematic evaluation of 12 open-weight LLMs reveals significant limitations in their ability to provide reliable causal judgments for structural discovery.
Researchers tested LLMs across six causal graphs using various prompting and confidence-scoring strategies. The study found that while LLMs can generate causal plausibility, their reliability as direct-edge classifiers is inconsistent, highlighting the need for better calibration before using them as automated priors in causal discovery workflows.