Elicit is built around a specific and disciplined activity, which is systematic review, and that constraint is what makes it good. Rather than offering to answer research questions, it helps you assemble a set of papers against criteria, extract comparable information from each, and see the shape of the evidence. That is a workflow with an existing methodology behind it, and building software that respects the methodology instead of replacing it with vibes is the correct decision. The honesty is what sets it apart.
The material states plainly where extraction is unreliable, which fields it covers less well, and where you must verify against the source. Most tools in this space assert accuracy and let users discover the limits at their own expense. Elicit tells you in advance, and the guidance around verification is framed as part of the workflow rather than as a disclaimer. Showing which part of a paper a given extracted value came from is the practical expression of that philosophy and it is the single feature I most wish competitors would copy, because it makes checking cheap enough that people actually do it.
Coverage is uneven and the material is upfront about it. Empirical work with structured methods sections, clear populations and reported effects extracts well. Theoretical work, qualitative research, humanities scholarship and anything where the argument is the contribution rather than the data extract poorly, because there is no field to pull. Users in those areas will find the tool much less useful than the marketing of the category generally implies, and I appreciate that Elicit says so rather than pretending otherwise.
Credits shape behaviour here as elsewhere. Systematic review is inherently exploratory in its early stages, you refine criteria by trying them, and a pricing model that charges for exploration discourages the iteration that produces a good protocol. That is a structural tension between the business model and the methodology, and the material does not acknowledge it. The assumption of methodological literacy is the gap I would most want closed.
The tool executes a systematic review workflow competently and does not teach it. Inclusion and exclusion criteria, handling of study quality, awareness of publication bias, what a synthesis can and cannot conclude. A user without that training gets an efficient path to a confident looking evidence table that means less than they think, and the risk is greater precisely because the output looks so professional. There is similarly little on the front end problem, which is that most bad reviews fail at question formulation rather than at execution.
A vague question produces a coherent looking answer to nothing in particular, and no amount of extraction accuracy repairs that. Three point seven. Easily the most methodologically serious tool in this category and the most honest about its own limits, held back by uneven coverage it declares openly and by an assumption that the user arrives already knowing how to do the thing the software is helping with.