Workshop at NeurIPS 2026
AI & Science Evolution or Extinction?
Call for papers
The AISciK workshop explores various aspects of scientific integrity throughout AI integration. Submissions should consider the AI system – its role in and impact on scientific practice – as the object of study. Papers that instead use an AI system as an instrument to produce a scientific result are out of scope. We ask submissions to treat measurement as something to be examined rather than assumed. When a benchmark, a task completion, or a study of scientific practice is offered as evidence, the question of what it actually measures is itself a research question. We are interested in how this question of measurement can be addressed from diverse perspectives including fundamental science, philosophy of science, science and technology studies, AI interpretability and alignment, sociology, and more.
Technical contributions are welcome and should be directed at meta-scientific questions — safety, alignment, interpretability, and evaluation as they bear on scientific knowledge production. Work that uses AI to understand how AI affects scientific research is in scope. Work that uses AI to advance a scientific result is not.
Not in scope
- Demonstrations of AI accelerating discovery in a scientific domain.
- New AI4Science methods or benchmarks whose contribution is capability measurement. Critiques and meta-evaluations of such benchmarks are in scope.
- Works that focus on AI’s impact in a way that is non-specific to science.
Tracks
Tracks describe the form of a contribution, not the discipline it comes from. For example, conceptual, historical, and qualitative work belongs in Research when it reports a completed investigation; the Perspectives track is for arguments, whatever methods the author works with. If a submission could plausibly sit in two tracks, pick the one you find most fitting — we will move it if the reviewers think it belongs elsewhere.
As this is the inaugural AISciK workshop, we have provided a short list of full-length papers we believe illustrate the types of work that would fit well in each track, either in terms of subject or methodological approach.
-
Research
Completed or substantially advanced investigations. This includes empirical studies, theoretical and formal work, philosophical and conceptual analysis, historical case studies, and qualitative research. What the track asks for is a result: something established, argued through, or measured, with the work to support it.
Submissions in this track are reviewed on soundness, significance to the workshop’s questions, and whether the methods used are appropriate to the claim being made. Submissions are assessed against the standards of the tradition they work in. Null and negative results are welcome.
Example Works
- Hao, Xu, Li, and Evans, Artificial intelligence tools expand scientists’ impact but contract science’s focus (Nature, 2026)
- Lee et al., The impact of generative AI on critical thinking (CHI, 2025)
- Kusumegi et al., Scientific production in the era of large language models (Science, 2025)
- Ayonrinde and Jaburi, A Mathematical Philosophy of Explanations in Mechanistic Interpretability (2025)
-
Datasets and Evaluations
Datasets, evaluation frameworks, observational studies, and critiques of existing measurement instruments. Contributions include:
- Benchmarks and evaluations designed to measure the impact of AI on scientific practice, or to target epistemic properties rather than task success
- Observational, ethnographic, and longitudinal datasets from scientific research settings
- Meta-evaluations and reproducible critiques of existing scientific benchmarks and agent evaluations
This track is reviewed on validity, documentation, and what the instrument can actually support a claim about. Artifacts should be accessible to other researchers; state where and under what terms.
New AI4Science benchmarks whose contribution is capability measurement are not in scope.
Example Works
- Alaa et al., Medical large language model benchmarks should prioritize construct validity (2025)
- Bean et al., Measuring what matters: construct validity in large language model benchmarks (2025)
- Reuel et al., BetterBench: assessing AI benchmarks, uncovering issues, and establishing best practices (NeurIPS, 2024)
- Singh, AI in Science (Data & Society, 2026)
-
Perspectives
Arguments about what should be done, what is at stake, or how a question should be framed. Position papers, disciplinary interventions, and contestable claims about the direction of the field.
Reviewed on the clarity of the position, the quality of the argument, and serious engagement with the strongest opposing views rather than convenient ones. Literature reviews and surveys without an argument of their own are not a fit for this track. A submission that could not be disagreed with is not making a claim. Submissions to this track should include an Alternative Views section in the main body of the paper that describes and addresses one or more viable (not strawmen) positions that are opposed to the paper’s position.
Example Works
- Hogg and Villar, Is machine learning good or bad for the natural sciences? (2024)
- Narayanan and Kapoor, Why an overreliance on AI-driven modelling is bad for science (Nature, 2025)
- Tang et al., Risks of AI scientists: prioritizing safeguarding over autonomy (Nature Communications, 2025)
- Barman et al., Large physics models: towards a collaborative approach with large language models and foundation models (EPJ C, 2025)
- Sullivan, Understanding from machine learning models (BJPS, 2022)
- Messeri and Crockett, Artificial intelligence and illusions of understanding in scientific research (Nature, 2024)
Submission guidelines
Format
Submissions are either 4 pages or 8 pages of main text, excluding references and appendices. The 4-page option is for concise contributions and extended abstracts; the 8-page option is for more substantial work. Neither length is preferred, and reviewers are not instructed to expect more from a longer submission.
References and appendices are unlimited. Appendices may be included after the bibliography, but reviewers are only obliged to read the main text, which must stand on its own.
Submit a single PDF in English, with any appendices included in the same file.
Template
In the spirit of cross-disciplinary inclusivity, we will allow submissions to be PDFs prepared using either the NeurIPS 2026 LaTeX style or another system if necessary. If using another system, please match the formatting of the NeurIPS 2026 LaTeX style. We have included the PDF version of the formatting if you are unfamiliar with LaTeX.
Please change the footnote to Submitted to/Accepted at/Published in the AISciK Workshop (NeurIPS 2026).
Anonymity
Novel contributions will be reviewed under double-blind conditions, requiring authors to anonymize the submission, including links and supplementary material, and cite your own prior work in the third person. Dual submission and previously published works are welcome. In terms of review, previously published works will be reviewed single-blind.
OpenReview
All submissions go through the AISciK OpenReview portal. If you do not have an institutional email address, account approval can take up to two weeks — create your account well before the deadline.
Submissions may be revised any number of times before the deadline. Revision is not permitted during review.
Disclosure of AI Use
Submissions must declare how AI systems were used in preparing the work, via the disclosure field on the OpenReview submission form. We expect that many authors will have used these systems; thus, disclosure will not be held against a submission.
Works substantially generated by AI systems are not eligible. The workshop is premised on human scientists deciding the future of science with AI, and a submission is a contribution to that decision. Responsibility for the content rests with the human authors, who are accountable for everything submitted under their names, including fabricated results or references.
Presentation and Attendance
Accepted papers are presented as posters. Best paper awards will be selected from among accepted submissions.
The workshop is built around discussion, and we expect one author of each accepted paper to attend in person. If circumstances make this impossible, contact us at aiscik.workshop@gmail.com — we would rather hear from you than lose the contribution.
Policies
Archival Status
The workshop is non-archival. Acceptance is not a publication of record, and there are no proceedings. Authors remain free to submit the same work to a journal or conference afterwards without prejudice.
Accepted papers are posted on the workshop website and on OpenReview. Rejected submissions are not made public and are not deanonymized.
Dual Submission and Prior Work
We place no restrictions on where else a submission has been or is going. Work under review elsewhere, preprints, work presented without proceedings, and work already published are all welcome, provided submitting here complies with the other venue’s policy.
Previously published work is welcome for a reason: much of the relevant research on scientific integrity has appeared in venues the AI community does not read, and vice versa. Bringing that work in front of an audience it would not otherwise reach is part of what the workshop is for. Published work is reviewed single-blind and need not be anonymized.
Withdrawal
Authors may withdraw a submission at any time before notification. Withdrawn submissions are removed from consideration and are not made public.
Reviewing
All submissions will require at least one reciprocal reviewer, unless no authors are qualified, in which case an exception can be requested from the workshop organizers.
Review process
Submissions are assessed against the standards of the tradition they work in. A conceptual paper is not judged for lacking experiments, and an empirical paper is not judged for lacking a philosophical argument. Reviewers are assigned by track and by disciplinary fit, and the reviewer pool spans philosophy of science, science and technology studies, sociology, evaluation, AI safety, and the physical and life sciences. Previously published work is assessed on fit and on what it contributes to the workshop’s questions rather than on novelty.
Reviewers read the full main text and confirm having done so. Decisions are made by the organizers on the basis of the reviews; there is no rebuttal phase.
We are recruiting reviewers across all the disciplines the workshop draws on, and particularly from outside machine learning. If you would like to review, sign up below. Authors submitting to the workshop are encouraged to volunteer.
Use of AI for Reviewing
Submissions themselves must not be shared with AI systems, in whole or in part, under any circumstances. A submission under review is confidential material entrusted to the reviewer, and putting it into a model breaks that confidence regardless of the model’s data policy.
Reviews must be written by the reviewer. Responsibility for a review rests entirely with the person who signs it, including for anything a model contributed to it.