Georgetown UniversityAI, Analytics, and the Future of Work

About

Our Mission

The AI, Analytics, and the Future of Work Initiative addresses critical issues regarding the economic and social transformations brought by digital technology. Through research, convening, and education, the initiative informs business leaders and policy makers on how to devise solutions that help the most vulnerable members of society and foster the common good.

About this leaderboard : Research Objective, Method, and Disclaimer

Research objective

The AI Referee Paper Leaderboard is a real-time research experiment conducted by Georgetown University’s AI, Analytics, and the Future of Work Initiative. It examines whether frontier AI models can identify promising economics and finance research before later conference and publication outcomes are observed.

Assessments are recorded as working papers enter the FEN network. As outcomes become available, we will examine whether papers receiving stronger initial scores are more likely to be accepted at selective conferences or published in leading economics and finance journals. This prospective approach allows us to test whether AI assessments contain meaningful early signals rather than simply recognizing papers that are already prominent or published.

The project is not an endorsement of replacing peer reviewers, editors, or conference committees with AI. Its purpose is to document the real-time capabilities and limitations of frontier models when applied through a standardized, practical implementation. Whether these technologies should be used in consequential academic evaluation raises separate questions about accuracy, bias, accountability, confidentiality, governance, and human oversight.

How the leaderboard works

The current leaderboard uses Claude Opus 4.8 as an AI referee for recent SSRN finance and economics working papers. The model reviews each paper’s full extracted content against a fixed rubric covering:

Each paper receives an overall score from 0 to 100. Its accompanying report provides dimension-level scores, perceived strengths and weaknesses, suggested improvements, and the model’s assessment of its confidence. See the full methodology and scoring rubric here. The current leaderboard reports the top 100 papers at different frequencies: 1M, 3M, YTD, and 1Y.

Interpreting the results

Scores and reports are experimental signals, not definitive judgments or validated predictions. Each score is a snapshot tied to the manuscript version, model, rubric, and evaluation procedure used at the time of assessment. A score does not establish a paper’s accuracy, originality, importance, methodological soundness, publication potential, or an author’s ability. It should not serve as the sole basis for decisions about publication, conference participation, hiring, promotion, funding, or reputation.

AI assessments may be influenced by limitations and biases in the model or its training data, manuscript presentation, disciplinary conventions, and the wording of the prompt and rubric. Models may favor polished or familiar forms of research, undervalue highly novel or interdisciplinary work, overlook methodological problems, or generate inconsistent conclusions.

Reviewing the contents of a paper in entirety does not mean that every proof, calculation, citation, dataset, or piece of code has been independently verified. Model-generated confidence is the model’s characterization of its own assessment, not necessarily a statistically calibrated measure of accuracy. Human expertise, careful reading, replication where appropriate, and conventional peer review remain essential.

Tracking models and outcomes over time

Frontier AI models change rapidly. The leaderboard’s results reflect the model, manuscript, rubric, and implementation used at the time of each assessment. We will track model versions, assessment dates, and material changes to the evaluation process. Where feasible, we will also evaluate common sets of papers across successive model generations. Scores produced by different versions may not be directly comparable.

We will compare these contemporaneous assessments with later conference and publication outcomes to study:

Conference acceptance and journal publication are themselves imperfect indicators of research quality. They are used here as observable outcomes for studying AI capabilities, not as complete measures of scholarly value.

Contact

AI, Analytics, and the Future of Work Initiative
Georgetown University
futureofwork@georgetown.edu

Learn more at futureofwork.georgetown.edu.