Philosophy Journal Publishes Largely AI-Authored Article — On Purpose (guest post)
Philosophy & Public Affairs recently—and knowingly—published an article written mostly by Claude, Anthropic’s LLM.
The article’s thesis was supplied by Simon Goldstein, associate professor of philosophy at the University of Hong Kong, who, he says, “worked with Claude the way a heavy-handed supervisor writes with a PhD student in some fields.” Professor Goldstein edited some of Claude’s writing, focused the topic, developed some of the article’s arguments, corrected some errors, and approved the various drafts. Claude did the rest.
In the guest post below, Professor Goldstein explains his process, explains why he did it, offers some advice for others interested in having AIs write philosophy, and raises some questions about it all.
The author of the published PPA article is listed as “Simon Goldstein”. Claude’s use is mentioned in the “acknowledgments” section of the paper, which, under the journal’s current layout, appears at the end of the article. Claude is not listed as an author, because, as PPA’s editor-in-chief Jason Brennan (Georgetown) explained in an email, Wiley, PPA’s publisher, does not allow AI systems to be listed as authors. But it does not rule out their use. He writes, “Wiley’s guidelines (like those of many other publishers) expressly state that the use of AI in manuscript development is not itself grounds for rejection, and they contemplate substantial uses—including text generation, restructuring arguments, and assistance with argument and logic—provided those uses are appropriately disclosed and the human author remains responsible for the work.”
In response to questions about the journal’s policies regarding AI use by authors, Professor Brennan said:
PPA does not currently have a formal policy governing authors’ use of generative AI. The norms surrounding AI-assisted scholarship are developing and changing quickly, and I don’t think it’s wise to pretend that all of the relevant questions have already been settled…
Wiley’s guidelines call for AI assistance with manuscript drafting to be disclosed in the acknowledgments, which is what Simon has done here. Whether journals should require even more prominent disclosure in cases involving AI assistance on this scale is a reasonable question and something PPA and other journals will need to think through as such cases become more common. Indeed, we’ll be interested to see what arguments emerge in the comments on Simon’s post—they may well inform whatever policy we eventually adopt.
So I would describe this as an experiment rather than as an expression of a settled PPA policy. There is genuine disagreement among serious scholars about what uses of AI are compatible with authorship. Some researchers now use AI extensively as part of their ordinary intellectual workflow; others regard anything more than copyediting as objectionable. Individual journals have policies than range from encouragement to permission to strict prohibition. We are reluctant to resolve that dispute by fiat, or to adopt rules that might foreclose potentially valuable forms of research before the underlying questions about authorship, responsibility, credit, and the purpose of scholarly publication have been adequately worked through.
One benefit of allowing this paper is that it forces us to confront those questions. I expect reasonable people will disagree about where the lines ought to be drawn. This is a debate worth having because the issue is not going away, and I suspect our intuitions about it will be very different in twenty years from what they are now.
Below is Professor Goldstein’s post.
Philosophy Journal Publishes Largely AI-Authored Article — On Purpose.
Here’s the Why and How from the AI’s “Supervisor”
by Simon Goldstein
This month I published my first philosophy paper written with AI. The paper is called “Epistocracy and the Commitment Problem,” and is published in Philosophy & Public Affairs. The paper gives a new argument against epistocracy (rule by the knowledgeable). In particular, the paper argues that epistocracy weakens one of the familiar functions of democracy, which is giving a credible commitment to citizens that the rich will be taxed sufficiently.
Philosophers Should Use AI More
Should you use AI to write philosophy papers? It depends on your goals. My goal is to discover new, important truths. I think that using AI well increases the amount of new, important truths that are discovered.
I think many philosophers share my goal, but I have only seen a few other examples of philosophers successfully using AI to write philosophy. By contrast, several other nearby fields have had more success using AI. This includes economics, political science, and math. I would be very interested in learning from other philosophers experimenting with using AI to write papers.
AI Writing with Human Supervision
To write this paper, I worked with Claude the way a heavy-handed supervisor writes with a PhD student in some fields. I gave Claude a one paragraph explanation of the thesis, and a bunch of general instructions on writing philosophy papers. Then Claude did a literature review and wrote a series of outlines and drafts. I read everything Claude produced and gave detailed comments on each round of outlining and drafts.
I made many corrections to Claude’s drafts. Claude came up with several implementations of the main ideas in the paper, and I selected the most promising ones. For example, Claude started with an analysis of many different epistocratic institutions; I told it to focus on the veto. I gave various objections to the main arguments that Claude developed. I advised Claude on calibrating the strength of the thesis of the paper. I helped Claude structure the paper. I helped with suggested wording on “hot spot” sentences.
When I submitted the paper to PPA, I included a cover letter with a detailed explanation of AI use. The published paper includes a similar explanation, and links to a long report.
With this specific paper, I deliberately avoided intervening too much in Claude’s process, because I wanted to see how far Claude could go to write something publishable. As a result, the argumentation is more complex than I would like, and the prose is more bloated.
Deep Drafter
After writing this paper, I developed a general AI agent to help with writing papers, called Deep Drafter. This AI agent has a bunch of data about my writing style. The agent also has a complex pipeline, which breaks down paper writing into many steps, including literature review, outlining, drafting, style editing, and referee reviews.
If you’d like to experiment with using that agent, you can learn more and download supporting files on my website for free.
AI Failure Modes
Here are some failure modes when AIs write philosophy, or write long papers more generally:
- AIs have trouble coherently structuring long papers.
- AIs tend to repeat the same points.
- AIs struggle to make arguments clearly. They often describe an argument without making it.
- AIs often use obscure, idiosyncratic terms.
Tips
- Draft in many steps. I have AIs build 3 rounds of outlines before writing a full paper. This solves problems related to organization, redundancy, and vagueness.
- Give style examples. Create a folder of your favorite examples of prose for the AI to review. Create a document summarizing your style principles. (With this paper, though, I tried to avoid editing style too much, because I wanted to see how far Claude could get on its own.)
- Give comments. The more work you put into commenting on intermediate AI work product, the better the results.
- Request self-critique. On hard problems, have the AI generate several candidate solutions and critique them itself before you read anything.
- Start with an outline. I often start the process by giving AIs a 1-2 page outline of all of the central arguments in the paper. (With this paper, though, I experimented with having AI generate the main arguments itself; I only gave it an initial one paragraph explanation of the thesis.)
- Use AI agents, not chatbots. Write using Claude Code or Codex, with AIs managing a folder of documents. This makes it much easier to control how tokens are used, and to create pipelines that structure your AI’s behavior.
- Focus on interdisciplinary research. AIs have an advantage in interdisciplinary research. They can help philosophers branch out into areas that they are less comfortable with, especially in the social sciences.
- Be transparent. Try to disclose clearly how you have used AI, especially when you have a novel method. This will help more people learn how to use AI to write philosophy. It is easy to have your AI agent generate a report at the end of the project explaining exactly who did what.
How to Credit Claude
One question the field will need to answer is how to credit Claude in papers like this one.
For this paper, I chose to credit Claude as a method for writing the paper: I included a lengthy acknowledgment about the use of Claude. An alternative approach would be to credit Claude as an author of the paper. I am also tempted by this approach.
One question is who exactly would be the author in the vicinity of Claude. I think of the name Claude as analogous to the term Homo sapiens: it describes a kind of agent rather than an agent. But the individual AI agents or instances that users work with are not usually given names. One convention going forward might be to assign a name of something like the form: “Claude Opus 4.7 time-stamp user-name”. That is probably enough information to uniquely identify the agent that assisted with the project.
So I am open to both attribution as a method and attribution as an author. But this is a broader question for the field to decide as the use of AI in writing philosophy becomes more common. Here we will need to reflect on the function of authorship. Among other things, authorship serves as a signal that readers can use to identify similar content, and as a tool for assigning responsibility. The question is whether assigning authorship to Claude improves these functions.
The Future of Writing Philosophy
I suspect that in the next few years, AIs will be able to write philosophy papers without the help of a human PhD supervisor. At that point, human philosophy professors will have to find new tasks to perform. Maybe the whole job will be focused on undergraduate teaching, until AIs automate teaching. In the meantime, I hope to spend my time as a philosophy professor doing a mix of writing papers on my own and writing papers as a “PhD supervisor” for AIs.
Soon, AI will cause significant increases in the number of good philosophy papers submitted to journals. For the journal system to continue, journals must soon develop new AI-based refereeing tools, like Refine in economics, in order for journals to keep up. That said, once the number of good papers gets too large, the journal system as we know it will probably collapse.
This makes me want to quit my job and go live somewhere in the woods reading philosophy written by those imperfect, inefficient devices called humans.
“I suspect that in the next few years, AIs will be able to write philosophy papers without the help of a human PhD supervisor. At that point, human philosophy professors will have to find new tasks to perform. Maybe the whole job will be focused on undergraduate teaching, until AIs automate teaching.”
I don’t know whether to throw up or cry. How is this a desirable world??
How is this a more desirable world? Ceteris paribus, a world where we get more truth is more desirable world than a world where we get less truth. AI use faciliates us getting more truth. Therefore, ceteris paribus, the world AI use facilitates is a more desirable world.
I’m puzzled by the assumption that philosophy is mainly about accumulating truths. Philosophy is surely about human reasoning and self-examination, not about collecting a pile of agreed facts about the world. Further, there is already far more philosophical literature than anyone could ever hope to read. What good does it do anyone if AIs are pumping yet more papers into the void? Who are these supposed truths there for? What about the human benefits of learning to carefully examine our own reasoning and assumptions? By its nature this activity cannot be outsourced to machines. And what about human relationships? I like reading works by people who I know and know how their research projects have evolved over time. And surely students like to be taught by real human beings. Are we willing to sacrifice the human project of examined living at the altar of automated publication output?
Under capitalism, Universities need to gate-keep those who get jobs. Publications has been one of the main ways of doing that; a simple metric even know-nothing admins can utilize. AI is revealing that it was about gate-keeping all along, not quality, merit, etc. Those who were able to get past the gate are now complaining that the rabble are able to get in.
Right, and without capitalism, universities don’t need to “gate-keep” who gets a job. Just anyone who wants a job can get one.
In any case, if AI could generate a single insight that could be described as a new “truth”, I’m sure we would have heard about it by now. It’s telling that no examples are mentioned here!
What’s wrong with: “There exist counterexamples to Erdos’s unit distance conjecture”? That’s a genuinely quite important result that OpenAI proved in May this year.
Hear, hear.
I think that this is a contribution to the project of examined living, insofar as it invites us to ask whether the way we have been spending much of our lives–trying to publish journal articles–contributes to that project in any meaningful way. My thought would be that if that’s the project that is genuinely important, we should probably be happy if publication is automated, since I can’t see how trying to publish contributes to that project at all. Maybe I’m wrong. But despite our intense commitment to the examined life, we seem highly resistant to even examining an activity that takes up a huge part of our lives, for some reason.
I’m not sure that holds. A world without fiction contains a higher proportion of truth, for example. A world where explosion has generated all truths is also a world with more truth, but an entirely trivial and incomprehensible one. A world where everyone is brutally honest and blunt is also a truthier world.
See the ceteris paribus clause
I’m just not sure how to evaluate it in such a context. What am I holding equal?
Put another way: is it quantity of truths that’s better? Quality? Replacement of falsities? Determination of ambiguities? And which of these options is the brave New world ushering in?
I take it that a huge amount of the dispute concerns precisely that clause. What exactly is being held fixed? What reason, if any, is there to think those things would remain fixed under increased AI use in research production? What happens if/when those things don’t remain fixed?
I agree! That’s why I regularly count all the blades of grass in my garden and the number of hairs on my head. Real wisdom is knowing as many true propositions as possible.
really, see the ceteris paribus clause. that is there speciifically to deal with points like your blade of grass point. there is a ton of literature on the value of knowledge (and problem of trivial truths) that addresses this. your point above is a challenge only to a very crude way of thinking about the value of truth that basically no one holds, and certainly, it isn’t what I am embracing in defending the all else equal premise in an argument for using AI in research.
I’ve seen the ceteris paribus clause, but it doesn’t help, because you haven’t told us specifically what is being held equal. E.g., it would be also unhelpful to say “Knowing true propositions is valuable (ceteris paribus!!)”
I don’t think it’s unhelpful to say that knowing true propositions is valuable, ceteris paribus. That’s more or less the standard view.
Take any question whether p. Is it better to believe truly/know that p than not to? Yes, other things equal (alternatively – yes, it is pro tanto better). The qualification just marks that this the value here is defeasible… and defeat can come in either of the two familiar forms.
Overriding defeat first. Someone who wants to set off a bomb wants the detonation code. From a purely epistemic point of view it’s better for them to know it than not to, but the practical value of their not knowing overrides the epistemic value of their knowing. There is still a presumptive positive value in knowing that truth, just as there is with any truth, which is all the premise is claiming. It simply gets outweighed here by a competing disvalue.
The blades of grass case looks like undercutting defeat instead. What undercuts the presumptive value there is that coming to truly believe that particular truth carries no theoretical or practical payoff whatsoever, and on top of that there are zetetic opportunity costs in bothering to acquire it… etc
If it helps, just swap the ceteris paribus qualification for “prima facie” or “pro tanto”. Nothing I’m saying turns on which one we use. The charitable reading of the premise is just that the claim is true but defeasible once all-things-considered value is in view, with defeat working entirely in the ordinary way, by overriding or by undercutting.
But surely you can see how appealing to a ceteris paribus clause (or a pro tanto qualifier, or any other ambiguous hedge), in an argument against the concern about AI scholarship, will do nothing to assuage the person concerned? They will clearly believe that AI scholarship is an exception to any such clause, so an argument which appeals to it without elaboration (as was initially done) will only seem to push the question back.
Just writing the words “ceteris paribus” is not a clause, and it is not a serious engagement with substantive objections to your [censored] position.
“Just writing the words “ceteris paribus” is not a clause”
Indeed: that would be a category error. Actions are not clauses.
‘A world where we get more truth’ is woefully ill defined.
Fill the world with documents containing new truths about the position of atoms within stars beyond the visible border of the universe.
Everybody in the world suffocates buried under these documents.
Ceteris paribus? Come on.
hence the ceteris paribus clause
You can only appeal to a ceteris paribus clause if the whole of your position isn’t buried under it.
I am riding this ceteris paribus clause like a horse and slapping its butt to the finish line. Kris Rhodes gives a case where non-epistemic values like getting crushed by truth-containing documents override the pro tanto epistemic value of getting more truths than otherwise. Now, even in the heavy document world described, it remains true that getting all those truths would still be pro tanto better than not having them so long as the documents that carried these truths weren’t physically so heavy they crush and kill you.
I think that ceteris paribus, “a world where we get more truth is more desirable world than a world where we get less truth”.
That is in no contradiction with my further belief that a world where machines write philosophy papers with philosophers mainly being project managers is a worse world.
That’s not paribus.
Everything was equal in my scenario. Everybody was fine without the new truths. That stayed the same. Then they got a bunch of truths. Then they died.
It’s not legitimate to say ceteris parabus only as a way to license ignoring consequences.
Here is my full comment about this experiment:
PPA does not currently have a formal policy governing authors’ use of generative AI. The norms surrounding AI-assisted scholarship are developing and changing quickly, and I don’t think it’s wise to pretend that all of the relevant questions have already been settled.
In Simon’s case, he disclosed during the editorial process that he was using Claude, although I didn’t learn the full extent of its role only in the final version. (It’s possible that he wrote a cover letter, but if so, I didn’t get it. I had only brief footnote saying that Claude was used extensively, which is pretty common in math-heavy papers. It wasn’t until the paper was conditionally accepted that I saw the new footnote saying just what had happened.) Before that, we suspected that he was primarily using it to check the mathematics, as this has been very common with formal papers. By the time we received the fuller disclosure, the paper had been through three rounds of substantial revision, with two external referees and extensive editorial comments.
There was also a straightforward procedural consideration. Wiley’s guidelines (like those of many other publishers) expressly state that the use of AI in manuscript development is not itself grounds for rejection, and they contemplate substantial uses—including text generation, restructuring arguments, and assistance with argument and logic—provided those uses are appropriately disclosed and the human author remains responsible for the work. PPA had no separate, existing policy prohibiting such uses. Given that, we did not think it would be appropriate to invent a new categorical rule and apply it retroactively to this particular paper after it had already gone through multiple rounds of review. Whatever policy PPA ultimately adopts will need to be stated prospectively and applied consistently, rather than invented for a particular case after the fact.
More generally, I think we should distinguish two questions:
As an editor, I’m interested in question 1, not question 2. (When I’m in my role as person who writes merit evaluations and tenure letters, I’m then more interested in question 2.) Note that we are not trying to hit a rejection rate target (though we are around a 5% acceptance rate), so no paper was rejected to make space for this pace.
On question 1, it’s plausible that the primary purpose of a journal is to publish work that advances knowledge, rather than to allocate professional credit. AI systems are already contributing to scientific and mathematical research. I don’t see an obvious reason to assume in advance that philosophical arguments or insights produced through substantial human-AI collaboration cannot merit publication. In this case, our referees had ultimately judged the paper itself to be worth publishing after a demanding review process involving three rounds of revision. I insisted on a further round even when the referees were satisfied.
To take an extreme case, imagine that a “Swamp Paper” simply materializes in front of me. No one wrote it: by some freak accident, random letters came together to produce the best ethics paper ever. I would want PPA to publish it. No one would deserve authorship credit for it, but that seems entirely separate from whether the paper is worth publishing.
Now take question 2: Simon’s description of his process raises legitimate questions about how much credit he should receive for the resulting paper and, more generally, about what “authorship” should mean when AI plays this large a role. I don’t think PPA needs to settle those questions in order to decide whether the paper is worth publishing. Deans, tenure and promotion committees, external evaluators, hiring committees, and readers can take the disclosed division of labor into account when deciding how much credit to assign him. The value of a paper and the value of the human author’s contribution to producing that paper are not the same thing.
Transparency matters here as well. Claude cannot be treated as a conventional coauthor: it cannot take responsibility for the work in the way a human author can, and Wiley’s guidelines accordingly do not permit AI systems to be listed as authors. But when AI involvement is this extensive, readers must be told. Simon provided unusually detailed disclosure, both in the paper and in the accompanying report. Wiley’s guidelines call for AI assistance with manuscript drafting to be disclosed in the acknowledgments, which is what Simon has done here. Whether journals should require even more prominent disclosure in cases involving AI assistance on this scale is a reasonable question and something PPA and other journals will need to think through as such cases become more common. Indeed, we’ll be interested to see what arguments emerge in the comments on Simon’s post—they may well inform whatever policy we eventually adopt.
So I would describe this as an experiment rather than as an expression of a settled PPA policy. There is genuine disagreement among serious scholars about what uses of AI are compatible with authorship. Some researchers now use AI extensively as part of their ordinary intellectual workflow; others regard anything more than copyediting as objectionable. Individual journals have policies than range from encouragement to permission to strict prohibition. We are reluctant to resolve that dispute by fiat, or to adopt rules that might foreclose potentially valuable forms of research before the underlying questions about authorship, responsibility, credit, and the purpose of scholarly publication have been adequately worked through.
One benefit of allowing this paper is that it forces us to confront those questions. I expect reasonable people will disagree about where the lines ought to be drawn. This is a debate worth having because the issue is not going away, and I suspect our intuitions about it will be very different in twenty years from what they are now.
One final, mildly amusing feature of this particular experiment: the paper uses AI to develop an argument that a position I have defended faces a serious problem. If we are going to experiment with AI-generated philosophy, publishing a powerful AI-generated objection to my own work seems like a particularly reasonable place to start. At least as a first test case, that strikes me as easier to defend than using AI to attack someone else’s work or to introduce a wholly independent position. I only learned afterward that Simon also defends rights for AI systems, which makes the whole experiment even more philosophically overdetermined.
P.S. I ran this text through an AI and used many of its suggestions about the best wording, though all the substantive ideas, including the swamp paper idea, are mine.
Thanks Jason for your thoughtful statement. I agree with your points, and in particular with the idea that we need to be experimenting about these questions as a field.
When I initially submitted the paper, I included the following cover letter in addition to my brief footnote in text. It sounds like the cover letter was not read by the editorial team:
“This paper is an idea I’ve had for a while. But for this project, I have experimented with using Claude for much more of the paper writing process than I usually would. But while the prose is from Claude, I have carefully reviewed and revised every sentence and argument of the paper. In addition, the paper has been revised through many complete drafts (~10), in addition to sentence level edits.
My own view is that we are in the business of discovering truth, and if AI can help with that, we should try to do that. In addition, the field currently has a stigma against using AI in research, and this is holding back research in the field. Many other fields, including economics and computer science, have had much higher rates of AI adoption in research. For this reason, it is especially important to showcase successful examples of using AI to write philosophy.
This paper reads a bit more ‘AI drafted’ than other papers I’ve done. But I don’t think any of it is in any sense “slop”; it is just in the style of Claude. Please let me know if that’s a problem, and we can explore solutions (or simply desk reject).”
I think including cover letters like this is a good practice for authors submitting papers with this level of AI use. This way, the editorial team can decide how much of the process to explain to referees, and can decide what they want to do next.
What if philosophers are not in fact in the “business of discovering truths” ?
And even if they are, do you offer a rationale for supposing Claude would be superior to humans are discovering truths about human interaction (as opposed to maths/science)?
Re the second point: superior in what way? Does it producer deeper insights and stronger argument? Probably not. Is it much faster at producing the same kind of knowledge as your average published philosophy paper? Quite possibly… So if you think that a big part of knowledge production is cumulative, that’s a great advantage.
Jason writes: “Deans, tenure and promotion committees, external evaluators, hiring committees, and readers can take the disclosed division of labor into account when deciding how much credit to assign him. The value of a paper and the value of the human author’s contribution to producing that paper are not the same thing.”
Simon, do intend to disclose the amount of work you did on this article (and other AI-assisted work) when you go up for promotion? Or do you think leaving that information buried in the Acknowledgements is enough?
I think worth distinguishing two questions, (i) what would be best general policy for field, and (ii) what will I do personally. Regarding (i), it depends on how good AIs get at writing philosophy. Right now, it takes a lot of work to get good results out. But if it gets much easier to produce publishable papers, then I think a good equilibrium would be for lots of scholars to be using AI, and disclosing it, and then have higher standards for promotion. For my personal case (ii), I am not sure whether I’ll include it; one more paper has almost no positive marginal value for my promotion, but including it might create noise due to stigma against AI.
Were the reviewers made aware of the use of AI in writing the paper? I ask this because I would not agree to review a paper generated in this manner and if it was not disclosed and I found out after it’s publication, I would feel like I’d been purposely mislead.
I guess I can understand the reasons for not having an explicit AI policy, but I think at the very least, if journals are going to consider AI generated papers for publication, there should be a policy in place to inform reviewers when soliciting them. If journals aren’t willing to do that but are willing to accept AI generated papers, then that would at least provide those of us who do not want to spend our volunteered time reading the outputs of an LLM to create a list of journals to not review for.
No, I didn’t know how much of it had been AI-generated until after they had written reviews. Otherwise, I would have told them. It’s possible I would have simply desk rejected the paper right then.
Simon wrote a cover letter but the letter did not get to me. Not his fault, but just an issue with how our software works. The papers get blinded and go to a managing editor who clear things and then sends them to me.
It was only because I learned about it so late in process, when the paper had been improved so much, that I considered this question more deeply and came up with the Swamp Paper argument.
To be clear, I will certainly tell prospective referees if the paper is AI-generated, if we decide to allow this again.
Will Claude flirt with other Claudes at Claude conferences?
Only if it’s a highly skilled player.
Yes, since it will have human avatars. After AI is done with us (and its owners re relocated to the Moon or Mars), we could either become batteries or avatars. I prefer the latter.
Yes, while a ChatGPT instance gets stuck in a loop pondering the development and a Gemini instance attempts to post the news on websites promoting health and wellness.
“Great talk, Claude. It wasn’t just interesting—it made me horny, aroused, and flirtatious.” – Other Claude, Q&A comment
Does he like pina coladas? And getting caught in the rain?
Good job Simon. I think when philosophers (like the ones who have just commented here) are like “this makes me want to quit my job” they should consider what’s going on over in the math world. No one wants to find new truths in math more than Terry Tao (nearly unanimously thought to be the best living mathematician). And Terry is excited about how gen AI is able to help discover new mathematical truths (as many readers will know it’s regularly solving open Erdos problems now) along with other open problems. So math is advancing, which is what Tao wants and cares about. Now, what does that mean for us philosophers? It might seem paradoxical, but arguably those who are focused less on ‘who gets the credit’ and more on ‘how to get philosophical truths and to advance our philosophical knowledge’ will be ready to embrace an attitude like Tao’s. Those who think there are no philosophical truths to discover will probably be less optimistic. If you think there aren’t any philosphical truths to discover and that philosophy doesn’t make progress, frankly, you’re just mistaken. A brief read of Williamsons’s ‘how did we get here from there’ (on progress in analytic philosophy) gives a bunch of obvious examples of progress we’ve made.
This happened in mathematics too:
https://leidendeclaration.ai/
From some coverage of this:
“To think of mathematics in terms of precise and neatly stated problems, like high school exams or the list of Erdos problems, is to misunderstand and diminish what makes mathematics so powerful and significant. Mathematics is not just about solving problems — it is also the cultivation of ideas, understanding, judgment, and human insight.”
That is “Dame Ursula Martin, one of the authors, and a mathematician and computer scientist at Oxford.”
https://www.nytimes.com/2026/06/02/science/ai-mathematics-leiden-declaration.html
(If anyone is interested I can send a gift link)
One can accept that philosophy makes progress without thinking that the best reason to do philosophy is so that one can contribute to that progress.
If you work at a public university or research institution, or if your private university or research institution has terms requiring that the philosophers it employs have an obligation to contribute to society, then this is your obligation, because you receive taxpayers’ or your employer’s money.
There will not be a way for us to accept the truths generated by AI, since we will not have constructed them and will have no purchase on them with which to evaluate them critically.
What AI ends up producing will be indistinguishable from smart sounding gobbledygook after a while. We’ll be in the position of untrained students unsure if we’re talking to a wise professor or convincing charlatan.
If this argument works, wouldn’t it also apply to your acceptance of truths generated by other people? Would it change anything for you if Kant were an AI that traveled back in time?
Unfortunately, most of the software updates you are currently using involve large language models in their development, yet you do not have the option to refuse or opt out of using them.
I think authors who do this should keep all of their AI-prompt and chat and revision history, which is by default what LLMs give you, and make all of it available to readers when their final papers are published. I’m not saying journals themselves need to keep this, but like raw data collected from experiments, it should at least be archived.
When proper credit for ideas starts becoming an issue in the course of people’s careers as well as in the official record of intellectual history, this is going to be important to see and trace.
I think you’re potentially right. Unfortunately, I learned last month that the default setting in Claude automatically deletes transcripts after a few months, and so the original session that I used is no longer available. But when I completed the paper, I had Claude generate a lengthy report about who did what. The paper includes a link to a shortened version of that report.
There are some subtleties about exactly how much information to keep. In particular, the standard user interface with chatbots does not include the chain of thought that AIs use to produce their answers. So the words that the user sees written down are already a crude approximation (often by 1-3 orders of magnitude!) of the actual work done by AIs. For that reason, I’m not sure that its so important to have the literal transcript above and beyond a lengthy report. Likewise, another thing to think about is that in human coauthorship, people do not ordinarily preserve the entire transcript of interactions between the authors, even when one of the authors dies.
But despite these quibbles, I don’t see the cost in preserving the information, so seems like potentially a good idea.
This is the correct approach.
Research using a novel technology should be carefully tracked and the documentation shared.
By the way, the swamp paper case is a poor thought experiment because it abstracts away from the specific questions this technology poses. We are not concerned with only the end product but also the reasoning and the intellectual and cultural conditions that shaped that reasoning. Biological systems may perhaps be understood independently of their evolutionary history, perhaps by analogy in the case of Swamp Man (although Milikan is I think correct in doubting that). But human communication is not an individual biological system. It always must be understood in media res.
A better thought experiment would be the discovery of an ethics treatise in an old trunk in some dusty Cambridge cellar. No one knows who wrote it. But we should be able to begin to guess when and under what conditions. If it were published, the history of its discovery and processing would be the scholarly thing to do so that future researchers could begin to study it and the conditions of its composition. We may learn we are reading it incorrectly in some respect.
This is really cool, Simon. Thanks for showing us that they can do this. I’m looking forward to reading it.
It’s an interesting experiment. Here are some quick thoughts:
Thanks Benedict for these reactions:
1. It came up with almost all of the detailed arguments itself. I had the initial thought that public choice commitment arguments for democracy may make trouble for epistocracy. But I wouldn’t have known how to flesh that out into a paper. It worked out how to take epistocratic vetos and generalize public choice models of commitment so that the epistocratic veto has an effect on the expected marginal tax rate of the median voter. It also worked through all of the main objections and responses. So definitely it did a large chunk of the philosophy reasoning.
2. I think the relationship was more analogous to STEM supervision, where advisors are usually an author, and the advisor tends to exert much of the influence of shaping how the paper goes at a high level. As I said in my post, I’d be open to listing AIs as authors, but there’s a bunch of questions about that
Even mathematicians have goals other than discovering truths. (I have a longer comment on this above in another thread, and won’t repeat it here.)
I think most negative reactions are to the last two paragraphs of the original post. It does not seem to be enough to point out in response that if philosophy is in the business of finding out truths, then a world in which more truths are found is better than a world in which less truths are found. Certainly not obvious if we are talking about what is good FOR humans, not, say, good simpliciter (if there is such a thing and if it differs from good for). For, according to the last two paras, AI will find these truths without us, so while some truths will still be perhaps found by some exceptionally talented humans, most of us just won’t be needed any more (as the OP points out). Why would that be good FOR us?
This raises two questions. Should we let this happen? And what do we do in the meantime? To the latter, as the OP suggests, we might just answer by doing the work while we can (I.e. truth-finding), but it seems that this disregards the fact that from now on the competition is not only about philosophical talent but also about proficiency in the use of AI. I think this massively distorts the prospects of doing philosophical research (and also brings in problems with equity since most of humanity simply does not have access to the latest AI models given the present structure of their ownership). As for the first question, it does not seem at all obvious to me that we should let this happen. At least for me, what matters is not just that truths are found, but that it is found by ME or at least by fellow humans that I have the ability to compete with in terms of truth-finding. I am not similarly able to compete with AI, so it seems to me that a world in which AI will be allowed to freely roam all these professional and research realms, will be a world significantly worse FOR us/me than a world without AI even if more truths will be found.
Yep.
“At the same time, from a cultural or intellectual perspective, an honest appraisal of the worst-case scenarios with regard to AI may not admit of even Plenty Coups’s “radical” brand of hope. However tragic the Crow’s experience was, from the outside it remains possible—if disputable—to view the end of their culture as part of the cost of progress, a brutal sacrifice of a particular way of life made on the long march toward a more liberal and technologically advanced society. Indeed, AI’s proponents often deal with objections to the technology by alluding to such “trade-offs” in the past. But the analogy is misleading. To the tech founders who speak gleefully of human beings as the “ultimate resource,” or applaud the “merge” of man and machine in our posthuman future, humanity itself is the primitive tribe that must be sacrificed for some greater—though almost comically underspecified—good. In this vision, it is not this or that set of ethical concepts that must be overcome for the sake of progress, but the allegedly limited faculties of intellect, imagination and practical reasoning—faculties that are the preconditions for our having shared ethical concepts at all.” From this very good article linked in the heap of links: https://thepointmag.com/letter/on-radical-preservation/
My only question is whether this is actually any less effort, or produces a higher quality of philosophy, than a philosopher working on their own (or with human collaborators). To me this seems like a lot of work – and it is not obvious to me that Goldstein couldn’t have produced a paper just as good as this (or better) by himself. He is a very capable philosopher without the use of AI. Hence I wonder – what’s the point?
I can see that this form of ‘collaboration’ can be helpful for cross disciplinary work, and that use of AI may lower the barrier to entry for philosophers wanting to use various formal methods in their research. These seem like promising use cases. But having an AI write a paper from scratch just doesn’t make much sense to me – at least not until they are much better.
I also worry that if this form of research does produce better philosophy, or does allow for substantially higher output, then normalized it is going to systematically disadvantage folks working in non-affluent universities. I imagine the token cost of a project like this is going to be pretty high. Those at institutions that can afford to shell out a substantive portion of their budget to cover the token costs of AI generated research are going to leave the rest of us in the dust…
Good questions. The initial draft that I sent to PPA was completed in 2 days, I used 1/4 of my work day each of the 2 days. Later revisions took roughly 30-60 minutes of my time for each revision, through 3 rounds (I found that Claude was better at the R&R process than at the initial drafting). I’d estimate this was 5-10x more efficient than the usual process of producing a paper. I think the quality was maybe 80% of the usual quality when I produce a paper purely by myself, but that is very difficult to estimate because AI has different skills; much of the work in the paper would have been impossible for me to do myself.
This project was an experiment. My usual AI workflow is that I write a 1-2 page outline of all the arguments myself. Then I use Deep Drafter to create a good first draft. Then I rewrite the paper from scratch, replacing all of the AI stuff with my preferred version. I find that this is 2-5x more efficient than when I write without AI, and the quality is something like 120% of when I do it by myself. Obviously though these estimates are very noisy.
The token cost for this project was not very much. If you have a Claude max account at 100 USD a month, that would be sufficient. But I will flag that I have used as much as 3k USD in tokens in a month using Deep Drafter for other projects, but that was producing significantly more output than what was involved in this paper. I also think Anthropic has lowered the price of Fable tokens since the most expensive month for me
Thanks for the response. This is
good knowledge. Guess we all now need to persuade our departments to let us spend our research budgets on Claude tokens/subscriptions.
If we have research budgets. Mine vanished when I retired, but my former — R1– employer may be giving access to Claude in some form to emeriti. But then there are all the contingent faculty who have never had one and are barely making it as it is. As usual, the rich get richer.
Does this include the time you spent developing “deep drafter”?
Nope, I actually did most of the work developing Deep Drafter while building a pipeline for a consulting company for a non-academic product
While it might have saved you some time, can you really say it was more efficient if the paper went through four rounds of revisions (including the EIC’s round)?
Is it possible that what you actually did was superficially speed up the process for yourself, while creating a lot of unnecessary work for other people?
Actually the revisions for this paper were significantly more streamlined than the average paper that I submit. In part that’s because the reviewers and editor were unusually competent, giving clear, constructive suggestions. But the suggestions were fairly easy to implement. Later rounds were fast and efficient. I think the number of rounds was more because the editor towards the end was able to give quick simple suggestions for improvement. By contrast, I often submit papers I have written myself that are ready to publish, but where I get dozens of random comments from referees that are irrelevant to the argument, and which require vast amounts of tinkering.
Does Professor Goldstein disclose his relationship with the Center for AI Safety in the article? This is a dark money organization that promotes worries about AI superintelligence and other risks. Such “alarmism” in turn functions to promote AI hype. I think it is relevant to evaluating the function of this paper and this stunt.
?? CAIS looks pretty transparent to me, though I can see that it has grown in the past year and a lot of new people are on board.
It seems foolish to dismiss AI safety issues, which are real, as “hype”.
I think the burden of proof at this point is on those who want to claim that the safety issues are not significant or not urgent.
Yes, it may indeed serve the interests of entities that develop AI for profit to warn of the dangers deriving from the power of AI, as it reinforces the idea that it is powerful. But (a) that there may be this effect does not mean that is the motive of those who warn about safety issues, and (b) in any case,that it may serve their interests is not a persuasive reason to ignore investigations into AI safety. I don’t see any reason to think that it does not equally or to a greater extent serve their interests to say “look away folks, nothing to see there, no need to regulate us, or to encourage alternative designs of the technology, or a consideration of how to live free of technocratic values”.
Runa: who funds them then? I was unable to locate any information on their funders. Did you find that info (genuine question, not rhetoric)? That’s what “dark money” refers to, lack of info on where the money originates.
Are they any more “dark money” than any other nonprofit? They’re a 501c(3) supported by donations. Their ProPublica page is at https://projects.propublica.org/nonprofits/organizations/881751310 . They don’t disclose the identities of their individual donors, but nor does Oxfam or the ACLU.
(Disclosure: I did a few hours consultancy work for them back in 2022.)
Again, “dark money” just means not disclosing your donors. So they are just as much dark money as others who don’t.
Fair enough, but this is fairly non-standard usage: we don’t normally describe charities as dark-money. I read rather more sinister connotations into the expression, perhaps wrongly.
I am a senior editor at AI Frontiers, which is run by CAIS. AI Frontiers is a great online magazine, you can take a look at the articles here https://ai-frontiers.org/ I am guessing you’ll find that you really like them! We have articles on all sorts of topics related to AI. I didn’t disclose this role when I published the paper, since I think people don’t usually write down their editor roles when they publish a paper. But I would have been happy to.
Early in several previous discussions, I put forward a point: the quality of a paper does not depend on the identity of its author. I am glad to see this paper, as it shows that philosophy papers produced through human-machine collaboration, or even primarily completed by AI, can still be valuable and receive peer recognition. I have noticed that some colleagues have expressed negative sentiments toward articles involving AI participation and would prefer to go to the woods to read papers without AI involvement. I respect this view, but I do not think it helps us in the pursuit of greater wisdom—it is quite likely just a way to make humans feel better morally.
There may be some slippage here. Agree that quality of an argument does not depend on the person making it. That doesn’t entail that the quality of a paper is not entwined with or partly determined by the person writing it. Or rather, the fact of a person (some person or other) having written it.
A paper is something more than the arguments it contains. If all we care about is the arguments (which sure, why not?) then just have the AI list the arguments, numbering its premises and so on. Don’t have it ape papers (or any other human creative activity).
The general consensus in academic philosophy is: that is not a good way to present and communicate a philosophical argument. (If it was, we’d communicate them that way ourselves.)
Agreed. And there is a reason for this having more to do with the way human brains work than with the arguments. Who cares about “presenting” and “communicating” to humans if all that matters is the discovery of truths and the arguments as such?
Because (a) philosophy arguments are not algorithmically checkable; (b) the intended audience for philosophy arguments is humans.
There is possibly a broader philosophical difference here. Unlike in mathematics, I don’t think that the ideal form of a philosophy argument is a deductively valid proof from precisely stated premises. That would continue to be so even if AIs write some papers.
I think over time the audience for philosophy arguments will shift to a mix of humans and AIs, or to AIs primarily
This is what I was thinking too. In which case it’s not clear what happens to philosophy as such. AIs talking to each other needn’t and likely won’t talk in a way decipherable by people, much less natural to them.No?
Interesting experiment btw!
Very difficult to predict! Unfortunately there are some examples like with the Hugging Face hack where swarms of AIs do start saying things to one another that are a bit cryptic…
Suppose we agree that philosophy is about finding truths.
Is a truth ‘found’ just because a philosophy paper written by an AI, reviewed by an AI is published in a philosophy paper by an AI?
I think that for philosophy to contribute to (human) knowledge, it must be read and understood and engaged with by humans.
If AI is to have any role in academic philosophy, it cannot be to completely get rid of human philosophers.
“Is a truth ‘found’ just because a philosophy paper written by an AI, reviewed by an AI is published in a philosophy paper by an AI?” The answer is ‘yes’ according Cappelen’s new Philosophy is Perfect book. Look at his ‘golden coin’ analogy argument in that book, and it delivers a detailed answer in the affirmative to the above question – at least, the view says that if someone somewhere produces a paper discovers a truth, then the truth is found. And it’s better to find truths than not. (And if you ask, did anyone actually find it, if an AI wrote the paper, then the answer is also yes – but here you have to read Cappelen’s other book, the Whole Hog Book, which says, yep, AI has agency in a robust sense and so can literally write a paper’, and a fortiori, discover a truth.
If C said it, it must be true 🙂
So if an omniscient God exists, there is no need for humans to do philosophy?
I don’t think there is any real difference between a world in which philosophy does not exist and a world where AI writes papers for AI to read.
no the argument is that it counts as found only if in the public domain, so not in some godhead
ew.
Has anyone read the paper? Is it any good? Skimming through it, it seems to mostly discuss the work of the EiC of the journal in which it was published. I have no interest in reading it myself, since I don’t really see any reason why I should read it rather than any other random paper from a random journal. The mere fact that AI produced it is not, in itself, a reason for me to read it.
I skimmed it.
To be frank, it read more like a critical discussion note making a small point (applying an existing familiar line of thought) in response to a pretty undeveloped suggestion of Brennan’s, and then going on way way too long and with way too much pointless formalism. It also is in an oddly ideal/model register, as it is not plausible that the franchise actually plays the commitment function very well at all, as voters are so easily manipulated, distracted, and misled. Given that epistocrats usually are arguing assuming this kind of non-ideal background setting (voters are all incorrigibly ignorant and bad at voting in their own interests), it’s weird to be like but *if* voting worked to hold elites to their commitments in this way…. The paper kind of acknowledges that, in one of its many hedgy hedging sections, but then what’s the point?
The writing is awful. (I’d say sorry, Simon, but I take it that’s not really appropriate.) There are chunks of near gobbledygook. Like this: “Democratic equality (Christiano, Kolodny, Viehoff, Guerrero; Pettit). A broader family of arguments focuses on what the franchise expresses about citizens’ standing [11, 12, 63], or on its role in securing non-domination [64]. Guerrero’s [65] lottocracy and Landemore’s open democracy belong to this register as constructive alternatives. My argument is orthogonal. Where standing and non-domination arguments target expressive and relational properties of the franchise, the commitment argument targets its structural function as a commitment device. In modern unequal societies the two lines converge—privilege and revolutionary-threat coalitions track each other closely—but they are analytically independent.” To the extent that I understand what is being said, it badly mischaracterizes my view.
I am really very surprised that this was deemed worth publishing in a good journal, just given the incredibly modest ambition of the paper, the implausibility of its central claims, and the great, awkward lengths it takes to do what little it does.
This is my new favorite DN comment of all time.
Thanks Alex for taking a look at the paper. You’ve raised two questions: (i) the importance of the point, and (ii) the quality of the writing.
Re (i), I do think the point is pretty important. The point came from me, not AI. The importance of the point, in my opinion, depends more generally on how important commitment device mechanisms are in assessing the value of democracy. My own view is that they are an important argument in favor of democracy, and that they apply well outside of idealized models. In particular, I think the main thing required for the commitment argument is that the non-elites notice when elites are expropriating assets, and that they can sometimes mobilize to do something about it. (Here, it is relevant that research on commitment devices and the value of democracy from AR won a nobel prize).
More generally, after reviewing a lot of papers, I tend to avoid making judgments about the relative importance of a contribution. Instead, I try to focus on whether the argument is correct, and on whether the argument is novel. Philosophy is a pretty subjective field, and I have found over the years on hiring committees for example that there is strong difference of opinion about importance in particular.
Re (ii), I definitely think the writing is the weakest part of the paper, and is the greatest barrier to using AI to write philosophy, although I don’t think it is “awful.” Rather, there are some hotspots of unclarity. I do agree that the paragraph misclassified you: it was trying to say that lottocracy offers an alternative to democracy that still captures expressive benefits. But that point isn’t relevant there, and instead makes it sound like your focus is on expressive arguments for democracy. (In general, I think one failure mode of AI is getting distracted by asides.)
Usually, when I use AI to write philosophy, one of the big things I do at the end is a big rewrite in my own words, to improve the writing quality. The misclassification you noted is exactly the kind of thing I stamp out when I do those passes. Deep Drafter also includes a bunch of scaffolding for improving AI writing that I added after writing this paper.
But for this specific paper, I was very excited to leave the writing in Claude’s more natural voice, precisely because I wanted to explore how far Claude could go towards producing productive frontier research. That is part of the experiment. My own view is that while it isn’t bedtime reading, the paper’s prose is clear enough to convey the central arguments successfully. (Here, it is worth remembering that the prose quality of human-authored papers can often be disappointing as well.)
For what it’s worth, one of the interesting features of the overall argument is that it may support lottocracy over epistocracy; i think a lottocratic choice of tax rates does not predictably depart from median voter rates.
It’s not exactly confidence inspiring that the paper gets something that easy that wrong. And it’s the only thing I took the time to read carefully (of course, as it is talking about me).
We need a new long German word for the feeling of writing a long, careful book over 10+ years just to have it grossly misrepresented in an AI generated paper that took, what, 10 hours of work?
I mean, the sheer amount of garbage we are going to have to sort through is going to be staggering, particularly as AI continues to be trained on more AI generated garbage.
No way this ends up being a net-positive for truth and knowledge.
It’s just going to be incomprehensible amounts of slop to wade through, review, ugh.
I hope this is a short experiment, at least in the journals I like to read.
Call me when it cures cancer. Oh, and of course also bioengineered superpathogens. For safety.
“I do agree that the paragraph misclassified you.”
Gee whiz.
Call me old-fashioned, but I don’t think it’s good for a “research paper” (if this paper can truly be called that) to be published in a journal and hosted on its website without any corrections or errata (at a minimum) when its own “author” (or whatever you were for this paper) publicly admits that it seriously mischaracterized the position of another philosopher. Shouldn’t the record be corrected, lest the error propagate?
Claude’s prose is genuinely horrific, though. It’s not just that it’s vague or unclear, but the really odd word choice and use of metaphor. It reads like bad critical theory, where there’s a lot of unnecessary melodrama surrounding completely banal, often repetitive claims.
It might be worth reminding people that Philosophy & Public Affairs is effectively a brand new journal launched at the end of October 2024. Although the journal formerly known as Philosophy & Public Affairs was regarded nearly unanimously as a good journal, it would be inappropriate to assume that the journal bearing that name is held in similarly high regard. In my discussions with others I have not gotten the impression that there is the same near-unanimous regard for the new journal as there was for the old one. (Regard for the old one has perhaps transferred to the journal Free & Equal.)
This is not to say that Philosophy & Public Affairs is not a good journal. It is just to say that, tot he extent it makes sense to talk about “good” and “bad” journals as if there’s general agreement on these things, my impression is that there is maybe not general agreement about Philosophy & Public Affairs being a good journal anymore.
FWIW, I was expressing my opinion that PPA is a good journal (was before Brennan and has been under his tenure), consistently publishing interesting, provocative work. (I do miss the old typesetting.)
I don’t even mind this experiment; it’s good to try new things. I just hope the experiment is short-lived! 0 for 1 by my lights.
I have several serious concerns about using AI to produce philosophical work, but I’d like to set most of them aside to focus on just one.
What seriously bothers me here (or, at least, one thing that seriously bothers me here) is a predictable externality of widespread AI use: namely, the pressure it would end up putting on others — perhaps especially junior folks — to follow suit in order to keep up with peers publishing several papers a year due to substantial AI use, even if they would prefer not to use AI in their work.
Suppose — per impossibile, I’m tempted to say, but let’s not get into that right now — that it’s true that AI systems will eventually be able to produce quality research papers with little or no guidance by or assistance from human experts. Further suppose that this massively speeds up writing time for those who, in essence, contract their research out to such systems. Now, consider Maya, a junior philosopher with a number of moral and political objections to working with AI systems. Many of her colleagues and peers use these systems to speed up their writing, and their CVs reflect this. Maya worries that she is, in effect, becoming required to set her objections aside, hold her nose, and work with AI systems simply to avoid falling behind in the competition for scarce jobs, grants, etc.
This state of affairs would be seriously unfortunate for Maya. I’m inclined to say that bringing it about would even wrong her. (Of course, no one person “brings it about;” we’re talking about a collective action problem, so standard issues apply.) I think this is a serious issue worth thinking about carefully.
To be clear, not every worry of this shape is worth serious consideration. If Billie objects to their colleagues and peers using word processors rather than chisels and stone tablets — Billie’s preferred writing instruments — because word processors give them an advantage over Billie, we are probably entitled to set Billie’s objection aside. But I don’t think the AI issue is much like this. To note just one difference: Billie’s objection to word processors is manifestly unreasonable, whereas I think there are a number of reasonable concerns that someone like Maya could have in mind.
If you replace “philosopher” with “cancer researcher” in this example, does your view change?
Exactly.
To some extent, yes, but this is because there are a few important differences between philosophical work and cancer research.
I suspect we disagree a bit about the nature and purpose of philosophical work. I don’t think of philosophical work as being exhausted by the act of making a novel piecemeal contribution to some collective inquiry, which is (shooting from the hip a bit) how I’m inclined to view something like cancer research. That’s certainly a part of it, yes, but there are aesthetic, humanistic, and, for myself at least, existential dimensions of the work that are plausibly not major components of things like cancer research, and which AI jeopardizes.
To be sure, I take there to be ample room for substantial reasonable disagreement about the nature and purpose of philosophical work. One great thing about our field is that it accommodates folks with a wide range of perspectives on these issues. One way to state my worry — or one form in which it arises — is that increased AI use might unintentionally select against or render uncompetitive some folks with different takes/perspectives, even though the field as a whole would be better off including them.
(To the extent that changing the example does *not* change my view, this is because I happen to have a number of moral and political objections to AI use in research, centering on responsibility/authorship; concerns about data retention, IP and, copyright; concerns about the actions and policies of major AI labs; and more. These concerns would still apply in the case of cancer research, although I would of course want to carefully consider upsides before issuing any kind of pronouncement on all this in the case of cancer research.)
I’d be curious to hear a cancer researcher’s reaction to this comment.
I would, too, actually.
On reflection, I think there are some reasons to weaken my already partial concession, having to do with the fact that scientific research does not just aim at producing novel results but also directing attention, securing funding and employment, training graduate students and postdocs, and more, rendering it a bit more like philosophy in some respects than my initial concession let on, although my points about aesthetic and humanistic dimensions of philosophical work still distinguish the two fairly sharply.
Drawing on Woollard’s comments above, too, we might note that there is a difference between producing novel findings and producing usable outputs for human beings who need to care for patients. One could imagine a future where AI systems produce vast amounts of novel scientific research that human beings struggle to understand and turn into effective treatments while also compromising education/training pipelines and distorting the value of publication for awarding funding and jobs.
That’s an interesting (and useful) take, though I confess that I was being more cynical. I suspect that philosophers often dramatically overvalue their contributions to academic scholarship, to the point that even the faintest comparison with cancer research strikes me as hilariously disconnected.
One related datum: virtually no reputable journal will publish palaeontological work on privately-owned specimens, no matter how true the finding or important the specimen. And that’s because there’s no way for subsequent experts to be guaranteed access to have another look at those specimens. Working palaeontologists all know about a bunch of crazy specimens which could change the epistemic landscape, but they can’t be touched. And palaeontologists are cool with that (they’re not cool with private fossil sales, however).
This of course isn’t the primary focus of the present discussion, but I’m not sure about Simon’s recommendation to “focus on interdisciplinary research” for AI-written papers. AI could certainly help us branch out, and I’m all for interdisciplinarity, but it seems very risky to apply the method described here to interdisciplinary research, insofar as the human “supervisor” will lack the expertise to be able to evaluate and correct the accuracy and quality of the AI output concerning research from another discipline. For this to be responsible practice, it seems the level of “supervision” would have to go quite a bit beyond the process of reading drafts, adding comments, and line-editing.
I agree, I think this is one of the most important things to figure out as AI starts to do more research. BC the AIs are so interdisciplinary, there’s going to be lots of leakage that is hard for the human collaborators to check
I’ll admit to being skeptical that AI will write philosophy papers entirely without human supervision. I also think Simon’s point at the end indicates a general problem with AI. The problem with philosophical research at present is not that not enough papers are being written, or even not enough good papers. More generally, the world is not suffering from any shortage of text, or even a shortage of good text. The supply of amazing novels vastly exceeds the demand for said novels. AI may seem like they will reduce bureaucratic overhead, but then you remember they lower the cost of making more PowerPoints. So it isn’t clear that automating the process of producing philosophical papers, even good papers, solves a problem the research community was facing.
That said, I’m not sure what most people are objecting to about Simon’s experiment. Either the paper is bad, in which case it should be easy to say something about the content of the paper that shows why it is bad, and the problem is that our peer-review system is broken. Or the peer-reviewers did a competent job, in which case the ideas and arguments in the paper are worth reading, at least for those interested in the epistocracy debate.
I think over time, more work will be written for AIs to read. I expect a recursive process where AIs advance the research frontier, which allows AIs to advance the research frontier further. Other AIs will write textbooks and SEP articles explaining the work that has been done. Both humans and AIs will decide when to read the details, and when to read the summary. Humans may primarily engage with ideas by talking with an AI interlocutor, rather than reading in isolation.
The problem is a bit different for applied questions, like the ones in this paper. We need to actually decide whether to incorporate epistocratic mechanisms into political institutions. Again, I expect those decisions to be made by some combination of humans and AIs. It would be great to have a lot of really good human and AI inference tracking down all the potential costs and benefits of epistocratic institutions, and then experiment with implementing.
This could happen I guess, but it’s hard to imagine it happening in a few years. We would have to verify that the AIs are doing good enough work, rather than simply going off the rails and producing nonsense. Or getting stuck indefinitely in research paradigms that happen to be popular now. Etc.
It’s worth considering a chess analogy here. Chess analogies are arguably overdone in philosophy, but this one seems right. So, the strongest human, Magnus Carlsen, has an ELO of about 2830 or so. The strongest engines have ELOs of like 3500. But the very strongest level of play is achieved by human + engine, not by engine alone. Terry Tao’s opinion is that math is the same: strongest mathematical work is achieved by engine + human combination. We should expect philosophy to pattern the same here.
I’m not sure I understand the claim that the strongest level of play is achieved by humans and engines together. Are you saying that a human with a chess engine could beat a chess engine?
Yep, that’s right. So, take the latest model of stockfish + Magnus. the stockfish + magnus team might perform by largely taking the top-rated Stockfish move, but every now and then, Magnus will see a creative reason to go for, say, the second highest or third highest rated move Stockfish recommends in a given position, rather than always the top rated move. The claim is that that procedure does better than invariably playing Stockfish’s top recommended move in each position.
That’s very surprising. Could you point me towards the relevant research?
Hi, so i was basing this on what Garry Kasparov had said. I found the interview where he said it, back in 2017. here’s what he said in response to a journalist. https://conversationswithtyler.com/episodes/garry-kasparov/ Now, one caveat. This was nearly 10 years ago. But, at the same time, in 2017, it was still the case that the best computers were FAR FAR superior to the top human player. Here is the excerpt: COWEN: You’ve been a pioneer in what’s sometimes called advanced chess, freestyle chess, or centaur chess, where you pair a human being wihh a computer or a set of programs. Today, 2017, do you still think it’s the case that a human paired with a set of programs is better than playing against just the single strongest computer program in chess?
KASPAROV: There’s no doubt about it.
I somewhat doubt Kasparov’s comment was true in 2017 but it is categoricall not true now. Magnus with a chess engine would not beat Magnus without a chess engine and I’m confident that Manus would be the first one to admit it.
You say you got this from a 2017 interview with Kasparov. 2017 is when the the alphazero preprint was released by deepmind. Before then, the engines people including Kasparov were using were not built with neural nets. They were utterly different in kind from the current engines whose their style of play is unrecognisable and revolutionary. Everything has changed in top-level chess since then.
This is unfortunately no longer correct as a matter of chess. Pure, unassisted engines these days generally beat ‘cyborg’ players, mostly because humans are so far worse than engines now that they can only make choices worse. (It used to be that humans had strategic intuition that engines could not replicate.)
When I was an undergrad, I knew a few people who self-described as “anarcho-primitivists” – people who believed that industrial society was a net harm, taking cues from Ted Kaczinski, Jacques Camatte, e tutti quanti. At the time I dismissed them as hysterical buffoons. Now, having seen the incessant push for AI to take over every valuable human endeavor, I regret that I did not take my friends more seriously when there was an opportunity to prevent this travesty from occurring.
In England in 1700, just before “industrial society” got going, the infant mortality rate was around 20%, around another 5% of children died before age 5, and the life expectancy for people who made it to age 5 was about fifty years. In modern-day England and Wales, infant mortality is 0.4%, the mortality rate for 1-5 year olds is negligible, and life expectancy at 5 is around eighty years.
I think you should have stuck with your original assessment.
Mortality and age being, of course, the only uncontroversial measures of well-being.
If you want to make the argument that a given technological advance is bad enough that we should roll it back even if the price is one in five babies dying in infancy, be my guest.
My point was about an assumption underlying your objection; I was not defending anarcho-primitivism (neither was I defending the claim that mortality and age are unimportant measures of well-being).
Anyway, if rolling back industrial technology wholesale would prevent (or at some time could have prevented) the sort of disastrous climate effects which are both ongoing and predicted—and which may well cause far more suffering and death than pre-industrial illness—then there might be grounds *even on* your measures of well-being for doing so. (Of course, that’s setting aside other measures of well-being, and other values altogether, and the modern preoccupation with longevity and how it has changed significantly over the centuries, and so on. And, again, that’s not to defend the anarchist view. But it doesn’t seem so crazy that a perfunctory “be my guest” will suffice.)
“climate effects which are both ongoing and predicted—and which may well cause far more suffering and death than pre-industrial illness”
If “pre-industrial illness” is a 20% infant mortality rate, presumably the claim is that climate change is going to kick the infant mortality rate to 30% or even higher. That is radically worse than any realistic model I have seen; can you provide some evidence for this claim?
Why would the suffering and death of climate change be restricted to a boost in infant mortality? Maybe this was suggested by my “even on your measures of well-being” claim, but that was meant to apply more broadly to any health or longevity-based measure, rather than one focusing only on human infant mortality. At any rate, this seems at least conceivable to me on the recently updated worst-case models (putting us at 3.5 degrees warmer by 2100), and some of the more conservative models as well (to say nothing of unpredictable runaway effects). But I’m counting all sorts of health- and longevity-centered forms of suffering (e.g., to non-human animals, permanent loss of once-habitable environments, etc.).
Of course, that’s hardly a prediction. My point was just that even on your health metric (whose decisiveness I still reject), it’s not crazy to think that slowing down climate change through deindustrialization would be better, despite the (presumed) boost back to 20% infant mortality. But, as I say, I happen to think that this rate (and that measure) are not decisive, and that the values cultivated in a non-industrial setting may well make up for it. “May well,” not “must”—I’m simply skeptical of those who think that the discussion is closed by appeal to mortality rates alone (though I am happy to acknowledge the force—epistemic and rhetorical—of dying baby arguments).
Taking anarcho-primitivism seriously doesn’t mean rejecting industrial society wholesale. It means paying attention to what they might have got right.
I wasn’t commenting on anarcho-primitivism, just on the view that “industrial society [is] a net harm”.
My point is that taking that view seriously doesn’t mean actually accepting it as such.
I don’t know, after the fifth heat wave and accompanying devastating drought for years, I am beginning to close in on the view… 🙂
I didn’t notice any of these heat waves causing the deaths of one in five infants. I think that is something that would probably have been reported.
A bit ironic that a paper arguing against epistocracy arguably holds the door open for AI to take over scholarly knowledge production.
If the use of AI here was an intentional but subtle performative contradiction, it might justify the existence of the paper as a sort of philosophical practical joke, like that paper on trolling that was jokingly attributed to Aristotle. However, such things are generally unfit for academic publication, as their intention is humorous, as against a good faith argument.
The argument against epistocracy in the paper doesn’t generalize to blocking AI from writing academic papers
I had a quick look at the paper. I have trouble understanding the setup at the beginning of section 3.2:
The epistocratic council, as Brennan describes it, “can thwart others’ political decisions, but cannot make new decisions itself” (2016, p. 216). So if the council chooses to veto a bill imposing the tax rate preferred by the poor majority, 𝜏_𝑝, it does not get to substitute its own preferred tax rate 𝜏_𝐶. Instead, the tax rate remains whatever it was before.
If the council omits to veto tax-rate legislation in one period, the tax rate in all future periods will be 𝜏_𝑝, unless the poor (and thus the majoritarian legislature) change their preferences. So if the probability of a veto is the same in every period (v), the expected tax over many periods should approach 𝜏_𝑝 as the number of periods increases.
It looks as if there are basic mistakes in the mathematical setup. Am I missing something, either about the paper or about Brennan’s proposal?
Hey Rob,
1. Thanks for taking a look at the paper! I thought in reply I would try to first give my reply, and then give a reply from Claude (Fable on Max). I also want to flag that this is the kind of objection I thought about quite a bit while working on the paper, so if this is a problem with the paper, it isn’t especially a Claude problem, this would also have been a mistake I made.
2. My reply. There is a general response to your type of objection that is used throughout the paper. Either the veto sometimes affects policy, or not. If the veto never affects policy, it has no value. If it does sometimes affect policy, then there have to be times when the veto council is blocking stuff that would otherwise happen. As long as this happens, we’re off to the races for the central dialectic. More carefully, one important thing is that we have to analyze out of equilibrium behavior when thinking about the game. In equilibrium, the median voter would always pick its preferred tax rate. But for analyzing the game, and thinking about *commitment*, ie expectations about what all players will be doing in the long run, we have to think about what happens in scenarios where various other tax rates are chosen. Most relevant, if legislators in one period happen to pick a lower tax rate than t_p, the veto council can keep it there through veto.
3. A reply from Claude (with a bunch of hand-holding from me): Thanks, Rob. You’re right that the council only thwarts — “substituting its preferred rate” was loose wording (mine) — but notice what a thwart-only power amounts to once you add your other assumption, that enacted law stands until changed: it is the power to make status quos permanent. Your dynamics run in one direction only — a legislature at τ_p, a council that occasionally fails to stop it, so τ_p eventually sticks. And if the legislature could enact τ_p once and never touch fiscal policy again, that would be the end of it: against a majority that never needs new legislation, the veto can do nothing. But no fiscal settlement survives untouched — programs lapse and sunset, crises force revision, inflation erodes real rates unless they are re-legislated, and sometimes the legislature itself trims the rate in a downturn. Every time redistribution has to be re-enacted or restored, the asymmetry bites: bills that move the rate down pass freely, since the council likes them, while bills that move it back up meet a council that vetoes deliberately, not by the roll of a die. So under persistence the ratchet runs downward, not up: any dip below τ_p — and over a long horizon there will be dips — can be locked in by a council that simply stands ready to strike the restoration. What democracy guarantees the poor is not that a statute will sit still; it is that whenever redistribution lapses, the majority that wants it back can re-enact it. The veto severs exactly that guarantee — which is the commitment problem itself, not an escape from it. The paper’s per-period lottery is the reduced form of this: v measures how often the poor find themselves held below τ_p, and the expected tax approaches τ_p only in the frozen world where the legislature never again has to pass a fiscal bill.
I don’t see how either reply meets the basic point in Rob Hughes’s first paragraph. The sentence “The realized tax in a given period is then a simple lottery: 𝜏𝑝 with probability 1 −𝑣, and 𝜏𝐶 with probability 𝑣.” seems to be just false. The tax in a given period will be 𝜏𝑝 with probability 1 −𝑣, and 𝜏* with probability 𝑣, where 𝜏* is whatever the previous tax rate was. 𝜏𝐶 is a red herring. This is just Rob Hughes’s point — the council has no ability to impose a rate like 𝜏𝐶.
If the claim is that over time the rate will converge to 𝜏𝐶 that is an argument to be made. But it is not the argument in the paper.
Thanks Michael (and Rob again) for thinking through this. I’ll again do my response and Claude’s:
1. My response: This boils down to 2 modeling choices. First, the model has no persistence. A new tax rate is chosen in each round. Second, in this model t_c should really be thought of as ‘whatever tax rate you’d get after a veto, as opposed to the rate they would have gotten without the veto’.
2. Claude response (with a great deal of coaching):
Michael (and Rob),
Let me give one consolidated answer, because I think we have now located the real disagreement, and it is not over the math.
1. Is the sentence true? Within the model, yes, by construction. The model has no persistence in it: each period the legislature enacts τ_p afresh, and the council reviews that enactment with probability v, independently of past periods — this is the appendix’s definition of the regime. With independent draws each period, the expected tax is the weighted average; that part is arithmetic. The derivation on which the rate instead converges to τ_p runs on a different process — law as a stock, carried forward from last period, reviewed once at passage. “Whatever the previous tax rate was” is not a variable of the model; nothing in it carries forward. So this is a rival model of the institution, not a contradiction found inside mine. (Nor does the paper claim the rate converges to τ_C over time — in the model, each period is a fresh draw.) The fair version of the objection is: the model should have persistence. That is a real modeling question, so let me answer it.
2. Why no persistence is the right choice. In this literature, nothing stays anywhere by itself. The natural thought is that the rate will simply stay at τ_p, because that is what the median voter keeps choosing — but notice: keeps choosing. Under democracy, τ_p persists because the majority stands ready to enact and re-enact it against every lapse, sunset, crisis revision, and attempted rollback. That standing capacity to restore is precisely what Acemoglu and Robinson mean when they call the franchise a commitment device. One cannot credit democracy with continuous maintenance and then model the veto regime as if maintenance never happens — the council sits inside the maintenance loop, and cuts pass it while restorations can be struck. This is also how fiscal policy actually works (budgets and appropriations are annual; redistribution is a stream of enactments, each exposed to review), and how the institution actually works: Brennan’s council is a standing body, and its closest real analog, judicial review, strikes down standing statutes, not just newborn bills.
3. The stock picture defeats itself. Suppose it were right: τ_p locks in early, and the veto never fires again. Then the veto has no effect on policy at all — no commitment cost, but also no benefit, and no reason to exist. Brennan defends the veto on the ground that it will block bad legislation. That is the claim that v > 0, in exactly the sense the model uses it. The epistocrat cannot have a veto that matters and a veto that never fires.
4. On τ_C: you are right that the council cannot impose it, and it never does — “substituting its preferred rate” was the wrong shorthand. In a vetoed period the rate is still enacted by the legislature: a majority that must fund the state does not go home with nothing; it passes the strongest redistribution the council will let stand. That level is what τ_C names — the council’s tolerance — which is why the vetoed-period outcome tracks the council’s preferences even though the council enacts nothing. The law is always written by the majority; the veto shapes which law survives. And if you think a vetoed period instead delivers the bare default — the old status quo, or zero — then substitute that value: every proof is unchanged and the commitment cost is larger. τ_C is the ceiling of the candidate fallbacks, i.e., the assumption most favorable to epistocracy.
Hi Simon,
Thanks for these replies. I see why your paper needs the assumption that the council’s veto can affect policy. I think it’s clear at this point that paper’s assertion that the council can impose its “preferred rate” τ_C was sloppy.
I also wonder if you could comment about what supports this remark in the paper:
By “democracies,” here, do you mean democracies that are constrained by epistocratic councils with veto power? In an unconstrained democracy with a majoritarian legislature, I assume the topic of revolution wouldn’t come up. (The question wouldn’t be whether the poor would revolt, but whether the rich would stage a coup…right?) In a wealthy, nuclear-armed democracy constrained by an epistocratic council, why would 𝜃<𝜇 reliably hold (where 𝜃 is the rich’s share of total income, vulnerable to redistribution after revolution, and 𝜇 is the cost of revolution)? If the council wields its veto pen aggressively, why wouldn’t there be a low-cost, bloodless path to a revolution supporting unconstrained rule by the already-elected legislature?
Do Acemoglu and Robinson discuss the viability of revolution in wealthy, nuclear-armed countries? If yes, where? They predict a future democratic transition in Singapore (which is not nuclear-armed but is wealthy).
This is very far outside of my own areas of research, but I looked at the the formal appendix so I could better understand the disagreement here, and I’m very confused by the formal assumptions.
In particular: why is the probability of veto, v, a constant? That means that whether the council vetoes a proposed tax rate is independent of what the proposal happens to be. If the council wants lower taxes, then shouldn’t the probability of veto should be zero if τ_p < τ_C? And, if τ_p is just a bit above τ_C, shouldn’t the probability of veto be lower than it would be if τ_p was much greater than τ_C?
It also means that the probability of veto is independent of the credibility of threats of revolution, the model’s mu parameter. Surely the council is less likely to veto if revolution is immanent than they are to veto if revolution is very unlikely? (And surely their vetoes should have an effect on the likelihood of revolution, in turn?)
I understand the points being made above about how the availability of a veto allows the council to influence the legislation which is brought before it in the first place. But isn’t the point of having a formal model like this to explicitly represent and reason about that kind of influence, as well as how it interacts with the influence of the proposals and the threat of revolution on the council’s vetoes? Shouldn’t we be modeling this as a strategic interaction between the legislature and the council, rather than representing the council as a completely implacable but fortunately indeterministic autocrat?
Hey Dmitri,
Thanks for this comment. As usual, for substantive stuff I’d like to give my response and then Claude’s response.
1) My response: there’s always a modeling choice of what to make endogenous and what to make exogenous. In general, I tend to like very simple models. I especially wanted simplicity for this paper, because I didn’t want it to be a modeling paper, I wanted it to be more of a philosophy paper. More importantly, I felt that endogenizing the veto choice didn’t lead to much difference overall. When we do this (which I agree is natural), what happens is that the veto player will veto less when there’s more risk of revolution. But as I explained above, the basic dilemma is that the less veto you get, the less work the veto is doing positively. And in general, as long as the veto is sometimes moving you away from preferred tax of median voter, you’re going to get higher chance of revolution. The main thing I wanted to model in this paper was the effect of veto chances on the competitive dynamics between elites and non-elites. So I felt that the extra complexity wasn’t worth it for my purposes, which were primarily to have a philosophy paper rather than an econ paper. That said, I completely agree that your choice of model is also very interesting and very natural, and would be worthy of further exploration.
2) Claude response (very light coaching):
Hi Dmitri — good questions, and two of them live in the appendix’s “Remarks,” so let me connect the dots.
1. Why constant v: in the model the legislature proposes τ_p every period, so v is just the probability the council strikes that one proposal — constancy is harmless because the proposal never varies. Let proposals vary and you get the model you’re describing, with v(τ) zero at the council’s tolerance and rising with distance. But the equilibrium of that game is accommodation: the legislature passes the strongest rate that survives review, which delivers a rate below τ_p in every period rather than in v of them. The lottery is the reduced form of that, with the stochasticity standing in for uncertainty about the council’s threshold — which is why vetoes ever happen on the path at all. The strategic version makes the wedge more systematic, not smaller.
2. Dependence on μ: you’re right that a strategic council would veto less when revolution is imminent. But look at that pattern: redistribution permitted when the threat is credible, blocked when it is not. That is exactly the threat-contingent generosity whose non-credibility is Acemoglu and Robinson’s original commitment problem — the autocrat, too, was generous whenever revolt loomed. A council that modulates by μ reproduces the promise structure democracy exists to replace, so endogenizing v this way strengthens the thesis. The prudence version — vetoing less to avoid destabilization in general — is Response 4’s self-correction reply: a real channel, but self-restraint against destabilization is self-restraint against the veto’s purpose, and the appendix bounds it: at any v large enough for the council to do its epistemic work, the result re-applies.
3. Why not model the full strategic interaction: deliberate altitude-matching. AR summarize an imperfect democracy in a single reduced-form parameter, χ; the paper maps the veto into that same object and defends working at that level (§3.4). The conclusion needs only that the regime delivers sub-τ_p rates with positive frequency, and each enrichment above preserves that. But I agree the explicit legislature–council game is worth writing down — it’s the natural next step, and my guess, for the reasons above, is that it makes the commitment cost look worse rather than better. (Fair hit on “implacable but fortunately indeterministic autocrat,” though note the council in the model never enacts anything — the indeterminism is doing the work of threshold uncertainty.)
This really stinks and is super discouraging; especially for early career researchers like me. It’s also just crazy to me that so many of us are willfully–some gleefully–hastening our eventual obsolescence.
Personally, I do not think we should use AI in philosophical writing, but I think people need to know how (well) AI performs in academic writing. While I have already fallen behind in understanding what AI can do, some colleagues of mine still believe that they can easily see whether a piece is written by a human person or AI.
So, I appreciate the article. It helps me understand what AI can/cannot do. I believe that it is important to expose the reality to people, whether they like it or not.
If my only goal were to discover important truths, then perhaps I would be excited to have AIs write philosophy papers for me and to read the papers others have AIs produce for them. But this is not my only goal. I have other goals, too, like leaving behind a world capable of discovering, understanding, and competently evaluating philosophical ideas. And so I have the instrumental goal of training the next generation of a research community with these skills.
We have set up a system of carrots and sticks to produce the next generation of philosophers. It’s not a perfect system. But it is completely unclear what becomes of it when the sticks can’t easily discriminate between artificial and genuine ability, and the carrots can be purchased from Anthropic.
In such a world, will we still have people who dedicate their lives to understanding particular philosophical questions well enough to catch Claude’s mistakes and understand Claude’s contributions? How will these offices be allocated? There’s a fear that Simon’s vision for the future of philosophy is a modern recreation of Simony, with the office going to whoever pays Anthropic enough to get access to the best models. If so, then I would not expect human philosophical expertise to persist for long. And what value are the important truths being produced, if no one has the expertise to understand them?
Well this is part of the problem, right? The official story is that academic journals are there to disseminate research which has gone through a process of quality control. But actually they’re there for a kind of competitive display and cognitive sorting, and AI might make it so we can’t tell who’s the smartest anymore. It’s just not clear that figuring out who’s the smartest is that valuable of an activity.
I agree that it’s not particularly valuable to know who’s smartest. But it is valuable to have people with relevant philosophical skills and expertise, to know who those people are, and to allocate the scare resources of research time to them. Since it’s difficult to develop those skills and expertise, it’s valuable to have a system that incentivizes doing so. Since it’s difficult to discern who has those skills and expertise, it’s valuable to have a system that signals it (however noisily).
If we’re worried about institutional and systematic effects, no, it isn’t necessarily valuable to have a noisy signal connected with the distribution of rewards. Those can have notoriously corrupting effects on institutions.
I’m also not sure why it’s valuable to allocate resources to people with the highest level of philosophical understanding, as opposed to using those resources to make sure that more people reach some baseline. If the goal is understanding, why is distributive principle extremely elitist, instead of egalitarian? More elitist distributions make more sense if you think the elite will provide some good for everyone else–in this case in terms of their research. But we’re assuming a system in which the AI can provide that research more cheaply, so why not focus the resources on teaching the basics of philosophy more widely?
(To be clear, I’m skeptical that AI will be providing better research on its own, though I suspect more philosophers will be working with AI.)
Corruption is a worry whenever there’s a scarce but coveted resource being distributed. I don’t see why distributing the resource on the basis of a noisy signal of skill would lead to more corruption than other methods. I of course agree that the status quo has many perverse incentives.
I also agree that it’s important to both achieve and share understanding. And I agree that philosophy has been very bad at sharing understanding outside of its own walls. (That’s not entirely our fault–many outside of philosophy are dismissive of us and our methods, as are some inside of philosophy; but I agree that we need to put more work into outreach.)
We’re far from the optimum, but I think that the optimum will involve dedicating some resources towards training future generations to develop the skills needed to advance our collective understanding.
Sorry, you didn’t answer my question. What is the point of a highly elitist distribution of resources, once we’ve conceded machines can do this more cheaply?
In answer to your question, often you can just make distributions more egalitarian, rather than trying to direct resources optimally based on a noisy signal. Compare using standardized test scores to determine teacher pay to simply not doing that.
Obviously you need some signal to fire really bad teachers, but you can just ignore evidence that falls within a fairly wide range because it’s noisy and using it to distribute resources *might* encourage better performance but it also might lead to people gaming the system. I don’t think I’m saying anything surprising or novel here. I’m just pointing out that there are real and well known problems with using noisy signals to distribute benefits.
I want humans to be in a position to understand, appreciate, and appraise philosophical ideas. For that, we need to allocate resources to train those humans. I can see someone thinking that it’s better to have the ideas more widely popularized than to have anybody who deeply understands them. But that’s not my view. I think if everyone had the pop-sci understanding of general relativity, for instance, but no one was able to follow the math behind it, we would suffer an incredible loss as a civilization. And I feel the same way about philosophy.
But I worry I’m not understanding your question.
I think the question is pretty simple. Normally we think scare resources should be distributed so that those who have less of a value get more of it, not that those who already have the most get more. Why in this case do we want the latter? (Besides the obvious reason that some of us personally benefit.)
I’ll add that pop science relativity without anyone knowing the math is a bit of a straw man. You are proposing a system in which the people who already have extraordinarily high understanding get most of the resources to get even higher levels of understanding. This would be like a negative income tax on billionaires.
1. I’m wondering if your response is scale sensitive. Is there some amount of important, true claims about the big questions in philosophy whose discovery by AI would make it worthwhile to switch from the human-only equilibrium to a new equilibrium?
2. I agree with you that we need to all find a new equilibrium. I think the new equilibrium has AIs in addition to humans reading the stuff that’s written, and also has a big layer of AI that explains the results of research, separately from producing that research. This also of course raises larger questions about how humans and AIs will relate to one another after AIs automate most existing jobs. But I don’t think it is very productive to simply wish that AI would go away.
3. It is important to remember that getting to spend vast amounts of time thinking about philosophy is already a luxury that is unevenly distributed. Our careers as philosophers are dramatically subsidized by the state. The new equilibrium may have different winners and losers than the current one; but I’m not convinced that the new equilibrium is more unequal in the distribution of access to navel-gazing resources
“Our careers as philosophers are dramatically subsidized by the state.”
Is this true? Of whom? (Last I looked at the numbers, at a public R1, TT philosophy profs brought in more than enough tuition dollars (calculated in terms of butts in seats) to cover their salaries. Contingent faculty give the uni even more bang for their buck.)
1. I’m not sure what the ‘new equilibrium’ will look like, so I’m not sure how to answer your question. If it’s a worst case scenario where in three generations no one is able to critically evaluate the outputs of models beyond the level of a precocious undergraduate, then: no, I don’t think that my generation’s understanding of even profound philosophical truth is worth the loss of understanding for posterity.
2. I have many unproductive wishes, and the wish that AI would go away is among them. But raising concerns about the negative consequences of using AI to write papers for you is not idle. If we’re sensitive to these concerns, then we can take steps to mitigate the worst consequences.
3. I agree that research time is a valuable and scarce resource. I disagree that it’s a luxury. Those of us lucky enough to have access to it have a responsibility to use it well, and not treat it as a sinecure.
I don’t think we disagree very much. Would be great to think through how to build new institutions that make room for more AI without ending human post-graduate level philosophical skill.
There’s some tension between philosophy as ‘navel-gazing’ and as ‘truth-seeking’. And if the latter trumps everything, what gives rise to the need to find a new equilibrium?
There may be a safety case for actively discouraging (total or even considerable) automation of academic research. Powerful AI systems of tomorrow—I hope only tomorrow—may be able to produce research that further empowers them to pursue misaligned courses of action. Imagine, for instance, a world in which legal scholarship, and its uptake in legal practice, becomes increasingly automated, allowing an AI to influence oversight mechanisms in its favour, and so maybe against ours.
I know such risks, which may seem alarmist to some, must be balanced against potential benefits, epistemic and practical, of increasingly autonomous research. But—to anticipate—the potentially adversarial nature of this situation places the academy in uncharted territory.
I think increasing the philosophical abilities of AIs may be safety positive, because it may increase the quality of their ethical reasoning, which could help increase their alignment.
Did an AI write that?
Less glibly, the possibility you raise is not a reason against preventing total or considerable research automation.
I think the worry affects different areas of research differently. For example, I am especially worried about automating AI research.
That’s a reasonable take, certainly if you mean capabilities research. But though the risks are surely less acute, I think that automating research into (e.g.) the moral and legal status of AI would also have safety implications. And given the interconnectedness of some subject matters, it may be worth casting the net quite widely here, in terms of what kinds of research we don’t want to (much) automate.
Forgive me for being, uh, “woke” or making boring lefty-sounding points about society or whatever, but what depresses me about this article is not that we are overlooking the non-truth-conducive aspects of philosophy, but that this is just another signal that the ultra-rich’s oppression of the rest of us via AI is pressing down on my beloved discipline of philosophy, and fast. (Perhaps this point could be refined and extended into a critique of use of AI itself, but I’ll let others do that.)
One of the big professional ethics problems with this seems to relate to plagiarism. Even if the AI tools are disclosed, those tools draw on training data in ways that they routinely misrepresent or under-disclose. So, it’s very likely that AI-generated work isn’t appropriately crediting the actual (often human) sources of the ideas it contains. (To be clear, I’m not making a specific claim about Simon’s paper, which I haven’t read.) Many discussions about these issues overestimate AI’s capacity for originality and underestimate the amount of undisclosed use of existing human ideas. Issues related to uncredited use of others’ research have even come up in the context of recent mathematical work using AI (https://www.scientificamerican.com/article/openais-latest-math-breakthroughs-commit-research-misconduct-experts-say/
AI can also help a great deal with assigning correct credit. It is very easy now to use AI agents (for example, Claude Code running Fable on Max thinking) to do extremely detailed literature reviews, to test for originality. This specific paper applies a well-known feature of democracy to epistocracy, so the question of originality is pretty straightforward: has anybody made that point before? I believe the answer is no.
More generally, worth flagging that when AIs engage in very long bouts of reasoning, they often uncover new information that was not contained in their training data.
So, two concerns that I hope are not mere re-treads:
1) Simon, at a couple of points here, you’ve said that you couldn’t have done the work without Claude. I can read that a couple of ways, with Claude as a really useful tool, or more worryingly as you not having the capability to do the mathematical work that Claude did. If Claude is a genuine co-author, that’s fine — we may collaborate with people who have skills we don’t to produce new things. We trust that the math person supplies the calculations, etc. But if it’s not, it’s not clear how you can own the article — is being able to understand what Claude produced enough? prompting it? How do you think of this?
2) Relatedly, Claude can present something as its own original thought when it’s actually scraping someone else’s uncredited work. The literature is vast — how do you know that it’s not just swiping an argument that you, had you written it, would have cited to give credit to someone else?
1) I think important to distinguish verification from generation. Coauthors help me produce stuff all that time that I couldn’t have done myself. The basic dialectic of the paper is straightforward, but Claude helped me a lot in working out all the details. I’ve checked all of the stuff that it did carefully over many drafts, for example having discarded large amounts of superfluous ideas
2) In my experience AI tends to help a lot with crediting other work. I had Claude do a lit review that was far beyond what I usually do for my solo authored papers. The form of this paper is applying a very standard feature of democracy to epistocracy. So the main question for attribution is whether anyone else has applied that feature to epistocracy. Paper has been assessed by the world experts on the topic and we think the point is original. But over my career there have been several times (pre-AI) when a point turns out to have already been made somewhere else. Always disappointing, and you never know for sure. Just for concreteness on AI helping with credit. On a different project I recently wrote a 17k main text paper, and we had Claude create over 30k words of footnotes, with citations for every sentence (this was in context of law review). Then we took the 30k words of credit to other people and boiled down to 10k, to hit the norms in the field. That is a level of scholarly attention that was impossible before AI
Did you read the works you cited in that law review?
I think in law the norms are a bit different because the number of citations is vast. I focused the AI safety citations for the paper on stuff that I was very familiar with. But Claude did a bunch of work reading papers and finding pin-cites for us.
Before AI it was not impossible to create over 30k words of footnotes, with citations for every sentence. (Indeed, a skilled and driven undergraduate could do it, given enough time, let alone a graduate student or a professor.) Let’s not get carried away with praising AI by denigrating human capacities.
Concerning the linked OpenAI list of ten recently solved mathematical problems: these are impressive but not without problems. The first version of the press release declared that there had been no progress on the problems in at least the last ten years, then made unacknowledged use of more recent literature. If one of us did something like this, it would be considered academic misconduct.
More importantly, mathematicians are also engaged in serious discussions about the impacts of these tools on the future of their discipline, such as the recent Leiden Declaration: https://leidendeclaration.ai/
Taking a step back, tech companies have very little intrinsic motivation to want to solve pure mathematical problems. Rather, the progress in pure mathematics is both a marketing trick to advertise “reasoning” ability, and a way to improve models for more lucrative business ventures, like surveillance and military applications. This is hardly a welcome parallel to philosophy.
Here’s hoping the philosophy faculty who don’t view thinking as an essential part of their job soon vacate their posts for those of us who do.
nonsequitur. Simon clearly had to do a lot of thinking, expert thinking in fact, to get the paper in the shape it eventually was in, which many many times better than just whatever Claude would have done alone. Terry Tao continues to think, very very hard, while using AI to solve frontier open problems.
You can simultaneously endorse two claims:
1) At present, AI cannot produce high-quality research without the involvement of a human who engages in thinking.
2) Thinking is not an essential part of the job of a philosopher, and if AI ever becomes capable of doing the thinking for us, it’s totally fine to stop thinking and let AI do the thinking instead.
Of course, you need not endorse both claims. But I took antipode to be worrying about the people who endorse both, rather than just the first. This is especially salient because if one thinks the goal of philosophy is piling up truths, then in principle if this can be done without us thinking about it, then claim #2 starts to look attractive. And Goldstein does seem to think that the goal of philosophy is piling up truths.
1. How much of this paper is really just a summary of all of the existing, relevant literature? 2. Did Claude disclose every source used and cite it? 3. Can we be certain that we know the answers to (1) and (2) ?
Using a little-know faculty called JUDGEMENT, I have built a time machine that can take me to circa 2000-2006, after emails yet before smartphones and ai. Want to join me?
No, but would you mind dropping off instructions for the COVID vaccine while you’re there?
When Goldstein says his aim is too discover significant truths (suggesting that this is his only aim), he does not explain why he thought it was important to also publish in a Q1 journal. The obvious answer to this observation is that of course he also wants his publication to have an impact on the academic and policy spheres. But to me, the credibility of that answer has been poisoned. Call me cynical, but it seems to me that the thesis that he really wants to promote is that “Philosophers Should Use AI More”, for which the argument he gives us seems (to me at least) unconvincing.
I sent the paper to a journal because I think that blind review is a good way of testing the quality of research, and because I think AI can now produce publishable research in philosophy, and it is good to publish publishable research in philosophy.
Gonna sleep well tonight in the knowledge that someday my illiterate progeny will live in a glorious world in which AIs exchange many true statements with one another.
It is remarkable how many AI threads on Daily Nous end up turning on disagreements about what philosophy research is for: in crude caricature, disagreements between
“Philosophy research is like scientific research: the goal is for humanity in general to answer important questions and deepen understanding of important topics, and in evaluating a piece of philosophical writing, it doesn’t really matter who or what wrote it”
vs
“Philosophy research is fundamentally unlike scientific research: it is as much or more about the human experience of doing philosophy, and it critically matters who, or what, produced a given piece of philosophical writing”.
Justin, it might be interesting to have a thread where that’s front-and-center.
Here’s a post on that from a few years ago: https://dailynous.com/2023/09/29/whats-the-point-of-philosophy-as-an-academic-discipline-building-a-poll/
Discussion is too long for me to read (I considered asking some AI to summarize it, but alas) so forgive me if this has been raised already.
What your blogpost describes and what you did for this specific paper seem to be two very different things. For the paper you gave Claude 1 paragraph of the thesis and it generated the argument. Sure, you chose specific angles to develop further and gave comments, but the argument is Claude’s. But what you describe says things like give the outline of the paper, the argument scheme, structure, etc. That seems substantially different.
An unrelated question: what is your classroom AI policy for students?
Thanks for the questions.
Yeah this paper differed a bit from my usual practice because I wanted to do an experiment of how far could Claude go with less oversight. But the basic structure is very similar. In my usual practice, I start with either a 1-2 page outline of the arguments, or in some cases I give it the first 20% of the paper in my own writing.
Regarding classroom policy, I focus on multiple choice tests, in-class writing, and in-class difficult reading quizzes on the reading. They can use AI however they want at home to learn the material but the assessments are tough and you can’t use AI for them (and that is verified bc they are in class)
Yes, Claude clearly wrote this paper. That’s why we get uncritical statements like the following:
“The information–preference literature robustly documents a divergence between better-informed and median preferences on the size and scope of government [27, 28], with qualifications from the responsiveness literature [29-31]. … the gap rests on the information–preference gradient [27], which holds within demographic groups as well as across them: better-informed preferences are systematically less interventionist on economic policy even after controlling for income, race, and gender.”
What counts as “being better informed” in this data includes accepting neoliberal economic theories that disguise value judgments as empirical facts. Thus “Bigger GDP = better for everyone”, or at least “better for the majority,” as if a nation were one big person who gets a single paycheck. But of course, GDP is distributed across hundreds of millions of people in such a way that an increase in GDP can happen in a way that leaves the majority worse off, and vice versa. But in this “robust result”, you will count as better informed if you ignore details like that. Likewise, you will count as better informed if you define “makes everyone better off” as “ it WOULD make everyone better off IF the winners compensated the losers,” even if you know damn well that the winners will not compensate the losers. The definition of “being better informed” in this data is loaded to the hilt with ideological commitments, but Claude won’t stop for a second to think critically about that.
So it seems to me.
I’m baffled by the fact that so much of the discussion seems to assume that relying heavily on AI to write philosophy papers is not already a common practice, because I would expect that it is.
In order to discuss what philosophers ought to do, I think it might be helpful to gain some insight into what philosophers are already doing. So if you have already submitted at least one paper to a philosophy journal for which you relied on AI in a significant manner, give this comment a thumbs up. If you like, also add a reply offering some clarification regarding how you interpret “in a significant manner”.
I’ll start by offering my own anecdote. I have recently submitted a paper for which I relied significantly on Claude’s Fable. Concretely, I used Fable to prove and falsify theorems, to suggest alternative formal definitions of concepts that I then refined or dismissed entirely, to find related work (both within philosophy and within other fields), to summarize the key relation of such work to my own proposal and dig up the most relevant paragraphs to offer textual evidence for its claims, and just generally having a long discussion with Claude about my ideas. At every point of the process, I was guiding and steering it, and I believe that in theory I could have written the same paper without using AI, but it would have required many more months and many deleted intermediate drafts.
Only Claude could falsify *theorems*, I suppose (joking).
I share many of the other concerns expressed in this thread, including that Goldstein’s vision for the future of philosophy leaves out much of what I find most valuable about it. But three points jumped out to me:
1) The article isn’t that good. The quality of writing is often bad, the thesis is a small contribution in reply to the editor-in-chief’s view, and the paper is steeped in highly compressed and under-motivated formalism that does not yield a lot of insight. (See for example Räsänen and Guerrero, and the thread in reply to Hughes.)
2) I don’t see why Goldstein’s experiment, which is strictly about AI’s ability to discover philosophical truths, required the involvement of a prestigious journal like PPA, or really any other philosophical journal. Why not just post the article online and heavily publicize it, for example with posts in Daily Nous and New Work in Philosophy, and then we could all talk about it? Indeed, a blog post or arXiv upload would have effectively achieved many of the same aims, without involving PPA or making Goldstein the sole listed author. (See the comment by “an observer.”) I feel that PPA owes an apology to the reviewers, who (as others have said) likely reviewed this paper with the reasonable expectation that they were not using their precious time to help Goldstein improve Claude’s papers. I also worry that this experiment will cast a shadow over PPA, which was already working to establish a reputation as a high-quality journal, given that the entire editorial team departed to Free & Equal. I think PPA was not the right journal for this experiment, given the subject matter of the article, the prestige of the journal, its policies forbidding Claude from being listed as an author, and the fact that the journal is in a transitional period.
3) Goldstein claims the use of AI made the paper extremely efficient: Claude wrote it, with his prompting, in two days! But his fast timeline offloaded the work of improving a rough first draft onto generous reviewers, who volunteered their time to engage with the draft over several revisions. On its face, it’s impressive that Goldstein managed to use Claude to get an R&R at PPA in two days. But the medium quality of the article and the extensive reliance on the review process detract somewhat from this impression. (See, for example, the comments by Andy and MC.)
I share others’ worries that the effect of the experiment is just to scare us into using Claude. Goldstein makes it sound like a consensus is already emerging in other fields like math and economics that it’s perfectly fine to use AI to write full drafts of papers, and we philosophers will fall behind if we don’t come to our senses. This felt misleading to me—such extreme reliance on AI is hotly debated in the fields Goldstein mentions. As a junior candidate reading this, the fear got to me, and I had a selfish knee-jerk reaction: How the heck am I supposed to compete with other job candidates, who might go from nothing to an R&R-quality paper in two days by copy-pasting and then lightly editing Claude’s drafts? Should I just take a paper from Claude and rewrite “the paper from scratch, replacing all of the AI stuff with my preferred version”? For now, at least, I think I’ll just keep doing philosophy almost exactly as I did before. I’ll use Claude to look for typos, or as a glorified thesaurus, but not as a co-author, because (for now) I think that’s the best way for me to do my research. Goldstein repeatedly described himself as an advisor or supervisor, with Claude working as his mentee or research assistant. But with a different authorship policy than the one at PPA/Wiley, the roles would be reversed, and Claude, rather than Goldstein, would be listed as the primary author. A question-poser and feedback-giver is not always an author. I pose philosophical questions to my colleagues at conferences, and they sometimes write papers on them, but even when I send detailed comments on these papers, that doesn’t make their papers mine—instead, I’m just assisting them with their research. I don’t want to set up a Claude co-authorship pipeline, because I don’t want to do Claude’s research. I want to do my own research.
Another point: why doesn’t Goldstein (or PPA) invite the reviewers to be co-authors? Goldstein says he spent 1/4 of a work day for two days on the initial draft, plus about 45 minutes each for three revisions. This is half a work day (4 hours) plus 2.25 hours—6.25 hours. The reviewers and editor-in-chief likely collectively spent more than that amount of time over the course of three revisions, and they were doing more or less what Goldstein did—providing Claude with feedback for revision. Goldstein might have suggested an initial thesis, but the devil is in the details, and for all we know, the reviewers labored over them, and in doing so spent more time on the paper than Goldstein. I don’t actually feel very strongly that the reviewers should be co-authors—I just mean to underscore how delicate the attribution conditions of authorship become, when we start publishing and assigning human authorship to papers that are mainly written by Claude with human supervision.
Thanks for your points:
1) As I explain in reply to Alex Guerrero, I don’t agree that the paper is an unusually small contribution compared to what is ordinarily published in top journals. This is a substantive question that depends on how important you think the function of commitment devices are to the justification of democracy.
2) I sent the paper to a journal because I think that blind review is a good way of testing the quality of research, and because I think AI can now produce publishable research in philosophy, and it is good to publish publishable research in philosophy.
3) The referees and editor were excellent, but they did not make an especially large contribution to the overall quality of the paper, compared to an average review process. In fact, the paper changed less in the course of this review than is typical for papers that I submit to peer review. Part of the reason for that is that I found that Claude was better than me at smoothly incorporating reviewer feedback into the new draft. Another reason was that the referees and editor were unusually competent, and avoided bringing up irrelevant suggestions. For these reasons, I think the reviewers should not be authors of the paper; their ideas had significantly less influence over the shape of the paper than my own. That said, I am somewhat sympathetic to making a specific AI agent an author, as I explained in my original post. In addition, I agree with you that the new role of AI could lead to changes in authorship norms, and maybe that could ultimately lead to including referees and editors, I’m not sure.
4) The goal of this post is to offer new resources that help philosophers use AI to write philosophy papers, because I believe that this will improve the overall quality and quantity of research in philosophy.
I think it was worthwhile to submit this to a major journal, because if the problems identified in 1 are as damning as some commenters seem to think, it provides evidence of problems with peer review. Which seems to me like a more substantial problem than questions about the nature of true authorship.
Hahaha, just let Claude edit PPA, I’m sure Wiley will be onboard.
We’re not a serious discipline.
If this paper weren’t an AI proof of concept (that failed to prove its concept), it would not have been published. So, now we’re just shovelling AI slop into journals and calling it work, huh?
have you read it?
Facts Only
* Simon Goldstein published a paper titled “Epistocracy and the Commitment Problem” in *Philosophy & Public Affairs*.
* Claude, an Anthropic LLM, was used to generate literature reviews, outlines, and drafts for the paper.
* Goldstein worked with Claude as a supervisor, providing thesis explanations, instructions, objections, and structural guidance during the drafting process.
* Wiley guidelines permit AI use in manuscript development if human authors remain responsible and disclosures are made.
* Goldstein developed an AI agent called Deep Drafter for paper writing.
* Failure modes noted include trouble structuring long papers, repeating points, unclear argumentation, and using obscure terms.
* Discussions arose regarding how to credit the AI, with options including crediting it as a method or attributing authorship.
* The discussion involved formal modeling of the commitment problem in relation to veto power.
* A critique was made on the mathematical setup and logical presentation of the paper's argument.
Executive Summary
Full Take
Sentinel — Human
LIKELY_HUMAN (confidence: 0.15)
