The SQARE initiative aims to establish a new quality assessment system for research articles in econometrics, offering an alternative to the quality signal that academic journals provide today. We are not building a journal and we will not publish papers. The arXiv already disseminates and archives econometrics research, and the journals already produce a signal that the profession knows how to read.
All of this serves a useful function. Our concern is with what that signal costs to produce, and with the fact that the system has not caught up with what the 21st century makes possible: an enormous amount of expert time goes into an old-fashioned and inefficient refereeing process, and a small group of editors makes high-stakes decisions. SQARE aims to arrive at a comparable signal more cheaply and more quickly; whether it does is the first thing we intend to measure. Because we do not publish papers, authors can submit to SQARE at the same time as they submit to a journal, rather than instead of it. Once SQARE is established, we hope that the SQARE evaluation, or one from a similar platform competing with it, will come to serve as an indicator of research quality equivalent to the role of journal publications today.
Where the project stands
We first set out this idea in May 2024. The original proposal, which 120 econometricians signed in support of its aims, remains available and unchanged, together with a fuller discussion of the background: see the original 2024 proposal.
In June 2025 the initiative received a grant from the UKRI Metascience programme, awarded through the ESRC: Fostering a Dynamic Academic Ecosystem: Innovative Platforms and Methodologies for Econometrics (UKRI reference APP47921), running to 2027 and held at the University of Oxford. The grant funds the design and construction of the platform, and the research needed to work out how the evaluation should be done.
Working through that research has changed the design. The central new idea is to assess quality by pairwise comparison rather than by absolute scores. Instead of asking a small number of referees to rate a paper against an implicit standard, we propose to ask a broad body of econometricians to make many simple comparisons between pairs of papers. Those pairwise comparisons are then aggregated into a ranking of all the papers in the pool, and ultimately into a SQARE score.
Before setting up the final SQARE platform, we will first run an experiment to test the pairwise-comparison and aggregation procedure. In the autumn of 2026 we will ask a large group of econometricians to evaluate a common set of arXiv papers, comparing pairwise comparison with conventional rating of each paper in isolation. What we want to know is how much expert time each method needs to produce a reliable ranking. We will publish what we find, whichever way the results fall, and the design of the platform will follow from the results rather than the other way round.
The planned SQARE evaluation process
A paper comes to SQARE, ordinarily one already posted on the arXiv. It then passes through two steps in turn.
A check of technical correctness. This asks a deliberately narrow question: are the proofs correct, the methods valid, the claims supported by the results? The check is automated and assisted by AI, using tools such as coarse.ink and refine.ink, which read a paper, follow its arguments and flag technical problems in minutes rather than months. It is a filter rather than a certificate: a paper that clears it goes forward to the editorial college, while a paper with flagged problems goes back to its authors with the specific issues, to be corrected and resubmitted. A contested flag is resolved by a human. Nobody is asked to reframe their contribution or reshape their paper to somebody else’s taste, because the check concerns correctness and nothing else. This step is fast and inexpensive, and its purpose is to reserve scarce expert attention for the judgement that machines cannot make.
Assessment by the editorial college. This is the substance of SQARE, and where it departs most sharply from the way a journal works. Members of the college are never asked to place a paper on a scale in isolation. They are shown two papers and asked a single question: which is the stronger contribution? Many such judgements, from many members, are aggregated into a ranking of all the papers in the pool. Because every comparison enters the same aggregation, the ranking is the work of the college as a whole, rather than of one editor and two or three referees. And because each member makes only a small number of comparisons, that work is spread thinly across many people.
Why work this way?
A comparison asks a judge only for something they can readily give — a direction — and never requires them to place a number on a scale whose meaning is private to themselves. Two entirely competent econometricians asked to rate the same paper out of ten may land far apart, not because they disagree about the paper but because they have calibrated the scale differently in their heads. Asking only which of two papers is better is meant to let that private scale drop out of the comparison. The idea is old — it goes back to Thurstone’s law of comparative judgment (1927) — and it has since been studied and applied at length in educational assessment.
The largest practical experience comes from exactly that field. No More Marking, the company behind much of that work, reports applying comparative judgement to nearly three million pieces of school writing and finds that it works well in practice: their judges find a pairwise comparison much easier to make than marking each piece case by case against a rubric. The largest published study, covering the writing of 55,599 primary pupils in England, concluded that the approach shows promise as a large-scale assessment method. School essays are of course not research papers, and comparing two econometrics papers asks more of a judge than comparing two essays — which is one of the reasons we are running our own experiment before committing the platform to the mechanism.
The ranking is turned into a SQARE score for each paper by a rule that will be public. We are still settling the exact form the score should take, and we will say more once the experiment has told us what the mechanism can support. Because SQARE has no pages to ration, nothing forces an artificial ceiling on how many papers can score well, and we intend the scoring rule to impose none. Nor need a score remain fixed forever: a paper can be returned to the comparison pool later — for instance once it has visibly shaped subsequent work — and its score recomputed by the same public rule.
Moving forward
Our immediate priority is the experiment described above. Beyond that, our next steps are to grow the group of supporters, to form the editorial college, to finish building the platform, and to run the first evaluations. EconBase, a companion platform for finding and navigating the econometrics research already on the arXiv, is being built alongside SQARE.
We begin with a single system for econometrics, but we value competition and experimentation, and we would be glad to see others, within econometrics and beyond, adopt this approach and improve on it.
Frank DiTraglia and Martin Weidner started SQARE in May 2024. This page describes where the project stands in July 2026.