The Alignment Journal: Organization, Personnel, and Scope

Cross posted to LessWrong. Please go there for comments and discussion.

The Alignment Journal is beginning to invite the authors of select papers to submit their work for review. If you are interested in participating as an action editor or a reviewer, make an account on our website; if you have a manuscript that you think would be a good fit for the Journal at this stage, email [email protected] to request an invitation to submit. Manuscripts under review will become visible on our homepage, and you will be able to nominate yourself as a reviewer of a specific paper that interests you. We plan to open up submissions to everyone sometime in October.

Here we announce the Journal’s inaugural senior editorial board, advisory board, staff, organizational structure, and initial scope. We welcome questions and proposed changes to help us refine the scope in the future.

Personnel

The Journal is run by its senior editorial board, which makes the Journal’s scholarly decisions, and two managing editors, who run operations and strategy. The advisory board offers high-level guidance without editorial responsibilities. A software lead and a head of operations support the team.

The senior editorial board (8 editors at launch, covering a range of alignment expertise) has authority over all of the Journal’s scholarly decisions and over approving and removing senior editors. The board focuses its attention on policy, standards, borderline cases, and appeals as it builds trust in the review process by fostering a strong culture of rigorous and impartial review. Senior editors may choose to handle the manuscripts themselves, but usually recruit an action editor (analogous to an area chair at a conference) who will be responsible for leading the review process for that submission, while the senior editors supervises. The full board of senior editors handle appeals and other contentious issues.

The Journal startup effort is being spearheaded by managing editors Jess Riedel and Dan MacKinlay, who have initial authority over operational and strategic decisions, preparing policy proposals for the senior editorial board but deferring to the senior editors on scholarly decisions. Editorial policy (such as refining the Journal scope, making acceptance decisions, and accepting new editors) will be decided by consensus of the full senior editorial board until a formal governance system is adopted. As the Journal takes on operations staff and eventually an editor-in-chief, this work should ultimately transfer to them.

The advisory board (7 advisors at launch) is composed of senior researchers who give high-level strategic advice to the managing and senior editors. Advisors do not normally handle papers or make editorial decisions, but they have a full window into the Journal’s operations and the editorial process. If they agree to take on the additional role of ombudsman, an advisory board member may hear complaints against senior editors and managing editors. The ombudsman role exists to ensure that complaints are heard by non-conflicted people.

The software lead, Yonatan Cale, has responsibility for creating and operating the Journal’s software stack, including its new editorial management system designed to support the Journal’s unique features.

The head of operations, Kristi Uustalu, manages communications, organization, scheduling, and finances, and interfaces with the Journal’s parent organization, Principles of Intelligence.

The Journal also receives occasional assistance from the staff of its sister project, ILIAD, and from Principles of Intelligence.

Advisory board

Senior editorial board

Managing editors

Legal structure

The_Alignment Journal _is beginning as a philanthropically funded project fiscally sponsored by Principles of Intelligence, the parent organization for the PIBBSS Fellowship. Significant operational support in launching the Journal is being provided by ILIAD, an applied mathematics AI alignment research and support organization that is also a project of Principles of Intelligence. The Journal expects to spin out as an independent nonprofit organization in the future.

Funding

We are grateful to the AI Safety Tactical Opportunities Fund, a pooled multi-donor fund, for providing our first year of funding, to the Survival and Flourishing Fund for supplemental support, and to NTT Research for in-kind support.

Scope

The Alignment Journal aims to clarify and address the challenges in designing a superintelligent AI that is aligned with human values. To serve that purpose, the Journal publishes work on the scientific foundations of understanding, predicting, and steering artificial intelligent systems.

We seek contributions that build theoretical or otherwise principled analyses of agency, robustness, incentives, generalization, interpretability, and long-term behavior across many learning and decision-making paradigms. We are especially interested in approaches that transcend contemporary architectures and potentially apply to future superintelligent systems that interact and self-improve.

A submission should make a claim the significance of which survives the deletion of any particular model, dataset, or benchmark, e.g., a theorem, an impossibility result, a formal framework, a protocol with an analyzed guarantee, or a conceptual argument precise enough to be wrong. Empirical work is welcome as validation or interrogation of such a claim.

Negative results are in scope and we would like more of them. This includes no-go and impossibility theorems, counterexamples to claimed guarantees, refutations of published results, and formalization attempts that failed for an articulable reason.

Alignment is an inherently interdisciplinary topic. To be accepted, work must substantively advance conceptual scaffolding or explanatory understanding of AI. Mere relevance or token connection is insufficient.

The topics we focus on are naturally influenced by the expertise of our editorial board, which will grow over time. We especially welcome work bridging alignment and these topics:

  • Causal inference
  • Computational mechanics and dynamical systems
  • Decision theory, game theory, and agent foundations
  • Information theory
  • Mechanism design and social choice theory
  • Moral philosophy, including metaethics
  • Multi-agent systems
  • Probability, statistics, and uncertainty quantification
  • Learning theory
  • Statistical physics and stochastic dynamics

as well as rigorous empirical studies of misalignment in frontier AI systems.

This list is suggestive rather than exhaustive, and we are open to other topics if they fit the aim of the Journal.

We currently are unlikely to review papers on the following topics, but may consider them in the future:

  • human-AI interaction
  • cognitive science and mathematical psychology
  • societal impacts, ethics, and governance of AI
  • security engineering and model-weight operational practice

Acceptance criteria

The Alignment Journal accepts articles for publication based primarily on three criteria:

  • Correctness: The claims in the article must be worthy of belief.
    • Mathematical computations must be accurate, proofs must be true, methodologies must be rigorous, and arguments must be sound.
    • Caveats and countervailing considerations must be presented in the manuscript with prominence appropriate to their seriousness (e.g., in the abstract).
    • Confidence in claims resting on empirical results is undermined insofar as those results are difficult to reproduce.
  • Insight: The article must advance our principled understanding of intelligent systems.
    • Improved benchmark performance per se does not generally imply insight, but the statistical behavior of models can of course provide important information about their internals.
    • Empirical evidence of improved performance on non-alignment capabilities is insufficient.
    • Scientific novelty in the traditional strict sense is not necessarily required: insightful translational and synthesis work will be accepted so long as it satisfies all criteria.
  • Value: The work should be the most valuable thing to read now for at least some alignment researchers.
    • It is not sufficient that a paper could be useful to someone at some time; the arXiv and journals/conferences with correctness-only criteria are more appropriate venues.
    • A result’s value to other researchers is greatly diminished to the extent it is difficult to independently reproduce and build upon from the methodology and materials provided.

Desk rejects

Work of any subject that fails the criteria above is declined without review. Common cases:

  • interpretability results reporting what was found in a model, without a formal claim about what such findings can establish
  • capability or propensity evaluations without an accompanying theoretical or conceptual advance
  • benchmark and dataset papers
  • unprincipled training techniques evaluated primarily by performance
  • attack and defence results that do not engage with theory
  • agenda and position essays lacking a strong technical or conceptual contribution
  • surveys without deep synthesis

Other publication factors

Preprint requirement

Each submission must be available (in non-anonymized form) on one of these standard preprint repositories: arXiv, SSRN, ECCC, PhilPapers, and PsyArXiv. Exceptions are made for paper types which are not permitted in the otherwise appropriate repository; in these cases, please make sure to include an explanation in the “notes for the editor” box while submitting. Email the editors with any questions.

Archival status and prior publication

The Journal is archival. The published articles constitute a permanent, citable version of record — the definitive, unchanging form of the work that is preserved indefinitely and treated as the canonical reference. The authors cannot publish the same work in another journal or conference. (Preprints on the arXiv or other preprint servers are encouraged of course.) That said, we adopt JMLR’s policy toward significant expansions of previously published work.

Specifically, we will consider submissions that have been published in a more limited form at workshops or conferences. In these cases, we expect the expansion to cite the prior work, go into much greater depth, and to extend the published results in a substantive way. In all cases, authors must (a) notify the editors about previous publication, including a link, at the time of submission and (b) explain the differences from their prior work. Examples of (possibly) acceptable ‘deltas’ beyond a conference paper include: new theoretical results, entirely new application domains, significant new insights and/or analyses. Examples of insufficient deltas include: adding proofs that were omitted from a conference paper; minor variations or extensions of previous experiments; adding extra background material or references. However, we ultimately leave the decision about whether a ‘delta’ is significant enough up to the individual reviewers.

Reproducibility

We strongly encourage authors of empirical and computational manuscripts to ensure their work is reproducible. In particular, we recommend they (1) include sufficient methodological detail, ideally a dedicated reproducibility section, and (2) upload data, code, and similar material to robust repositories. As described in the acceptance criteria, reviewers and editors should factor in the ease of reproducibility in their decisions. Where code, data, or other materials needed to reproduce or build on the work have not been shared, the reviewer abstract should alert the reader to it.


Credits and thanks

This post has been informed by gracious contribution and feedback from Gautam Kamath, Leon Lang, Konstantinos Voudouris, Geoffrey Irving, Edmund Lau, Yonatan Cale, David Udell, David Reinstein, Alexander Gietelink Oldenziel, Daniel Murfet, Zach Furman, and Marcus Hutter. All responsibility for errors resides with the authors.