1.0 Introduction
This is the second part of a series of responses to Rian Lobato. In Part 1, I addressed Rian’s claim that I am lying about my position on the evidence that moral realism is “extremely intuitive.” As I showed, this claim is completely baseless.
This article will focus on Rian’s second critical response to my claim that there’s no good evidence for the claim he made that “moral realism is almost always extremely intuitive across populations.” I’ll focus on the empirical case made in this post:
Dear Lance: there is a substantial experimental literature on folk metaethics showing that ordinary people often display objectivist-seeming moral judgments. Goodwin & Darley (2008), for instance, found moral claims to be regarded as substantially more objective than matters of taste or social convention and in some respects almost as objective as scientific claims. Nichols & Folds-Bennett (2003) found analogous objectivist tendencies in children and Beebe et al (2015) found broadly similar patterns of moral-objectivity judgments across Chinese, Polish and even Ecuadorian samples.
Now, none of that establishes that people are uniformly or universally moral realists. Wright et al., Sarkissian et al., and others provide evidence for substantial metaethical pluralism and context-sensitivity particularly when disagreement is framed across radically different cultures. And there is an additional methodological dispute over whether standard disagreement tasks actually measure metaethical commitments reliably (including, ofc, your own work with Moss - which I personally find pretty weak, since showing that standard probes are open to non-metaethical interpretations does not by itself establish that they systematically fail to track metaethical attitudes and at most it raises a validity concern, and revised later work has explicitly tried to control for alternative interpretations and still found high rates of objectivist responding, especially for clear cases of harm and injustice, but anyway). More recent work such as Sousa et al has in turn argued that some apparently relativist results may partly reflect those very measurement problems.
You may disagree (and you'll) with the evidence. But simply calling "unsubstantied" that objectivist or realist-seeming moral intuitions are widespread and robust enough to constitute a genuine empirical phenomenon across multiple studied populations is simply to lie.
Let’s summarize the claims:
There’s substantial evidence that ordinary people “display objectivist-seeming moral judgments.”
This doesn’t establish uniform or universal folk moral realism, because there is evidence of metaethical pluralism (i.e., individual differences in the extent to which people endorse realism or antirealism)
There are methodological problems with this research, including problems I’ve raised
My work with Moss is “pretty weak” because we only raise the possibility of alternative interpretations and subsequent work has corrected for this
While I may disagree with the evidence, calling it “unsubstantiated” is a lie
The second and third claims are true: there is evidence of metaethical pluralism and there are methodological problems with research in experimental metaethics. However, Rian’s first claim is ambiguous, and is either trivially true in a way I never denied or is false, depending on how the ambiguity is resolved. His fourth and fifth claims are both false. My research doesn’t merely show the possibility of alternative interpretations; I provide reasonably strong (if still tentative) evidence that unintended interpretations are extremely common and I conducted thematic analysis to categorize the particular forms these unintended interpretations took across well over a dozen studies. Furthermore, my multiple choice paradigms tested the extent to which participants can match the stimuli used in published studies to alternative characterizations of the same metaethical concepts. People did worse than chance. This is evidence of systematic unintended interpretation. Here, I’ll defend the following claims:
While there is substantial research showing that people give realist-seeming responses, the methods used in these studies are not valid. Thus, while it’s true that people may display what appear to be “objectivist-seeming” views, this doesn’t entail that they actually have realist views, or that the evidence that they do is good enough to justify believing that such views are common among nonphilosophers. I fully grant that studies report and appear to show that they have them; what I’m arguing is that these studies don’t provide good evidence that these appearances are accurate.
My work with Moss isn’t weak but this is largely irrelevant, since the Bush and Moss (2020) paper was an early article that preceded my dissertation. My dissertation provides far more comprehensive evidence against the notion that moral realism is common among nonphilosophers. Rian doesn’t provide a citation to clearly indicate whether he’s referring to my paper with Moss or to my dissertation, but the latter wasn’t coauthored and almost all the points and analyses there were my own, so referring to my work with Moss would most plausibly imply he’s referring to the preliminary article with Moss. If he’s not, that’s Rian’s fault for being unclear.
Claims that we have justification for or good evidence that moral realism is “extremely intuitive” across populations are not substantiated by available empirical data.
In short: there’s no good evidence moral realism is common among nonphilosophers, claims that it is are unsubstantiated, and my evidence for this is stronger than Rian thinks that it is.
2.0 An ambiguous thesis
Rian begins with an ambiguous remark:
Dear Lance: there is a substantial experimental literature on folk metaethics showing that ordinary people often display objectivist-seeming moral judgments.
What this sounds like is the modest claim that there’s lots of evidence that people display moral judgments that seem like they’re “objectivist.” What, exactly, does that mean? That they seem this way at first glance, or that they seem this way on reflection or given full consideration of those factors relevant to assessing the validity of the measures used and the best interpretation of these findings?
Rian is not clear. Something can superficially seem to be a certain way, but on closer examination not seem that way. Or it could seem that way even on closer examination, but still not actually be that way. It’s simply unclear what Rian is actually saying.
However, if Rian is merely saying that these studies superficially appear to provide lots of evidence that realism is common among nonphilosophers, this is definitely true and not something I’ve ever disputed. In fact, my arguments all rely on this assumption. What I argue is that despite it appearing as though we have lots of evidence people are moral realists, in fact we don’t have any good evidence this is true because these studies turn out to have massive problems.
I have never claimed there aren’t any studies purporting to show that lots of people are moral realists. Instead, I have consistently argued that while many studies claim to provide good evidence for this, that their authors are mistaken.
So if Rian were making this more modest claim, it wouldn’t make any sense. I don’t dispute this claim and never have. It would only make sense for Rian to be making a stronger claim roughly of the form that these studies provide good evidence that ordinary people “often” exhibit realist judgments as a matter of fact. This is precisely what I deny. I don’t think available evidence shows this, because I don’t think any of the studies used to show this rely on valid measures.
To try to clarify what I am claiming is ambiguous, and why it matters, suppose someone claimed the following:
Acme brand thermometers do not work. They don’t provide a valid measure of what the temperature is.
Now suppose someone came along and pointed out that Acme conducted lots of studies which purport to show that their thermometers do provide a valid measure of the temperature. They could say this:
There is substantial evidence showing that Acme brand thermometers provide accurate-seeming measurements of the temperature.
The problem here is the term seeming. Imagine we drop it:
There is substantial evidence showing that Acme brand thermometers provide accurate measurements of the temperature.
Notice how the moment we drop the term “seeming,” it becomes clear that the claim is that we have good evidence that Acme brand thermometers do, in fact, accurately measure the temperature. The inclusion of “seeming” throws a wrench in the spokes of any clear interpretation, because the claim that studies show “accurate-seeming” measurements is more consistent with the possibility that they only appear to at first glance but further analysis would reveal that they don’t, in fact, provide such accurate evidence. This would be consistent with what I’d say about the experimental metaethics literature: superficially, it seems like we have good evidence that nonphilosophers “often” endorse moral realism. I’m arguing that they don’t, in spite of how things seem.
But if “seeming” is taken to imply something stronger, as in evidence that isn’t merely superficial or “at first glance,” then “seeming” isn’t doing anything to assist and is at best superfluous. It shouldn’t be there.
Another problem is the use of the term “display.” This could likewise be leveraged to make a more modest claim. People could, in principle, “display” a realist judgment without actually making one. In other words, one could interpret the literature to show that when researchers attempt to assess whether nonphilosophers are moral realists, or make realist judgments, they “display” realist-seeming judgments only in the sense that their responses are in line with the operationalization researchers have employed for distinguishing realists from everyone else. If so, all Rian’s claim would consist of is the completely unobjectionable claim that we have lots of studies that are designed to assess whether people are moral realists, and that lots of people give the realist response in these studies. You will find that I explicitly draw this distinction in my discussions online about this research: I will emphasize that in certain studies most people give a realist response to the stimuli, which is distinct from them actually endorsing realism. You can choose a response that reflects realism without being a realist.
Again, if this is all Rian is saying, then of course I agree with that, and have explicitly said so publicly on many occasions. Indeed, my whole position turns on this being the case, since again, my position isn’t simply that it’s not the case that we have good evidence that nonphilosophers are moral realists, it’s that the evidence purporting to show that nonphilosophers are moral realists appears to show that they are, but fails because the measures used aren’t valid.
It would be unclear whether we disagreed about anything if Rian were making these more modest claims. However, because it would be inconsistent with Rian’s apparent position that we’re justified in endorsing his initial claim that moral realism is extremely intuitive across populations, I’ll interpret the “seeming” part of his remark as largely superfluous, and take his central thesis to be that we have substantial evidence that realism is intuitive among nonphilosophers. I think this is false.
As an aside, I’ve opted to use the terms “realism” and “realist” here for reasons recurring readers will be familiar with: these are more technical terms that do a better job of precisely conveying the relevant philosophical notion, whereas the terms objective/objectivist have too many colloquial uses misaligned with their narrower academic use.
3.0 Goodwin & Darley (2008)
Rian begins with the following claim:
Goodwin & Darley (2008), for instance, found moral claims to be regarded as substantially more objective than matters of taste or social convention and in some respects almost as objective as scientific claims.
This is a line from the abstract of the paper:
Experiment 1 showed that individuals tend to regard ethical statements as clearly more objective than social conventions and tastes, and almost as objective as scientific facts.
They are almost identical. Rian appears to have simply slightly reworded a sentence from the paper itself, and did so without attribution. This borders on plagiarism, but let’s set that aside. I think it is at least some evidence that Rian’s engagement with and depth of familiarity with this literature appear superficial; it seems Rian didn’t feel like putting the effort into describing the contents of this study in his own words. This is a bad sign.
But this isn’t the real issue. The real issue is that simply because a paper claims to show something, this doesn’t mean that it’s succeeded, or done a very good job of supporting the claim (and it’s worth emphasizing the claim is almost always that we have evidence for a conclusion; not decisive evidence that settles the matter). All Rian has done is tell us what Goodwin and Darley claim to have done. Did they succeed?
No. They did not. Now, I’m not one to reinvent the wheel, so let me first point out that Goodwin and Darley’s 2008 paper is considered a seminal work in the field among the few of us working in this area, and it has been subjected to extensive scrutiny. For instance, you can find an entire paper dedicated to critiquing their methods and conclusions from Pölzler (2017). Here’s the abstract:
Moral realists believe that there are objective moral truths. According to one of the most prominent arguments in favour of this view, ordinary people experience morality as realist-seeming, and we have therefore prima facie reason to believe that realism is true. Some proponents of this argument have claimed that the hypothesis that ordinary people experience morality as realist-seeming is supported by psychological research on folk metaethics. While most recent research has been thought to contradict this claim, four prominent earlier studies (by Goodwin and Darley, Wainryb et al., Nichols, and Nichols and Folds-Bennett) indeed seem to suggest a tendency towards realism. My aim in this paper is to provide a detailed internal critique of these four studies. I argue that, once interpreted properly, all of them turn out in line with recent research. They suggest that most ordinary people experience morality as “pluralist-” rather than realist-seeming, i.e., that ordinary people have the intuition that realism is true with regard to some moral issues, but variants of anti-realism are true with regard to others. This result means that moral realism may be less well justified than commonly assumed.
Pölzler argues that the best interpretation of their findings would be that they support pluralism, not realism. The claim that people regarded moral claims as “substantially more objective than matters of taste” and “in some respects almost as objective as scientific claims” is highly misleading. To be clear, I only endorse the claim that a surface-level reading of these studies is best interpreted in line with pluralism; this does not mean I necessarily regard the studies as valid measures that indicate pluralism.
There are two issues here: an interpretative issue and a measurement issue. Pölzler’s point here is that a proper interpretation of the results would be that they would (at best) provide evidence of pluralism, not a tendency towards realism, in contrast to what the authors of these studies claim about their own results. That these studies also rely on invalid measures is a position I take in addition to this; so I am not agreeing with any interpretation of the data such as “these studies provide valid measures that indicate pluralism.” I’m merely agreeing with Pölzler that a careful reading of the results of these studies shows that they don’t even appear to support the claim that participants in the sample lean towards realism in a less pluralist way. Responses are highly variable across conditions. If someone hated three out of seven of something, and liked four of them, this could average out to them liking members of the category if you averaged across a Likert scale. But it would be highly misleading to say that they had a positive outlook without explicitly acknowledging intrapersonal variation of this kind.
We’ll get to problems with the validity of the measures used, but first let’s talk about the problem with Goodwin and Darley’s interpretation of the data. Goodwin and Darley’s (2008) first method asks participants whether a given moral statement is true, false, or an opinion or attitude:
How would you regard the previous statement? Circle the number.
(1) True statement.
(2) False statement.
(3) An opinion or attitude. (pp. 1343-1344)
Goodwin and Darley interpret (1) and (2) as evidence that the participant is a realist, and (3) as evidence that they are less disposed towards realism. They also employed a second method in which participants were asked to consider a disagreement between themselves and another person who had previously taken the same survey. They were then told that this person disagreed with them, and were asked whether:
(1) The other person is surely mistaken.
(2) It is possible that neither you nor the other person is mistaken.
(3) It could be that you are mistaken, and the other person is correct.
(4) Other.
(1) and (3) are interpreted as realism, while (2) is interpreted as subjectivism.
They then combined these into an aggregate measure, where people were classed as most realist (giving a realist answer to both), intermediately realist (giving a realist answer to one and a nonrealist answer to the other), or least realist (nonrealist answer to both).
There is no gentle way to put it: these measures are terrible. The issue isn’t simply that these measures might not work, but that it’s unclear whether they could work, given the number of problems with each of them. But before even getting to issues with the validity of these measures, there is a serious problem with how Goodwin and Darley interpreted these findings, which Pölzler takes to be the biggest issue with them:
In what follows I will therefore focus on what I take to be the above experiments’ main problem, namely their inadequate metaethical interpretation of subjects’ responses.
As Pölzler points out:
Consider experiment 1. Goodwin and Darley assume that those who answer that a given moral sentence is (R1) “true” or (R2) “false” tend towards realism, and those who answer that the sentence is (R3) an “opinion or attitude” tend towards subjectivism. But both of these categorizations are inadequate. To begin with, R1 and R2 are not only consistent with realism, but also with all other variants of cognitivism, i.e., with subjectivism and error theory. Subjectivists believe that moral sentences are true or false depending on whether they correctly represent the subjective moral facts. Error theorists believe that all moral sentences are false. Furthermore, R3 only appeals to non-cognitivists, for only they believe that moral sentences cannot be assessed in terms of truth or falsity at all. By contrasting subjects who opted for R1 and R2 with those who opted for R3, Goodwin and Darley thus did not measure the prevalence of realism versus subjectivism, but rather of cognitivism versus non-cognitivism […]
I agree. Distinguishing people who think moral claims are true or false from those who don’t would at best measure whether people are cognitivists or noncognitivists, not whether they are realists or subjectivists/antirealists. In other words, this measure completely lacks basic face validity (i.e., the extent to which a method appears to measure the phenomenon of interest). Goodwin and Darley simply and straightforwardly misoperationalized the measure they were intending to capture. This would be the equivalent of intending to design a thermometer but accidentally creating a device that measures what time it is. This measure simply measures the wrong thing. Given this, if it provides any evidence at all about whether people are realists, it would do so only by accident, and in any case the onus would be fully on proponents of the study’s efficacy as a measure of realism to show that it measures realism in spite of its total lack of face validity. I don’t know that anyone has tried to do this, because it’s patently clear to anyone with a basic understanding of metaethics how hopeless this would be. Either way, this method has been abandoned as a measure of realism, presumably because everyone realized it doesn’t work. Rian appears totally oblivious to this, completely underestimating how serious the methodological problems with this study actually are.
It is worth pausing here to emphasize just how much of a problem this already is for Rian. Rian has uncritically presented what is arguably one of the worst studies on this topic in the entire literature, with apparently no appreciation for this fact. This serves as considerable evidence that Rian has at best a superficial understanding of the literature, and is hardly in a position to make confident pronouncements about how dumb or dishonest I am. To be clear, I am explicitly stating that Rian’s particular choice of study undermines his credibility as an informed and knowledgeable person about this literature, though that’s not saying much since I haven’t seen anything else from Rian that would suggest familiarity with the literature. Nobody working in this area would point to this study’s findings if they were trying to show that there was good enough evidence to justify the belief that realism was extremely intuitive. It was one of the earlier studies and has more flaws than almost any subsequent research.
But I digress. This is only the beginning of a tsunami of shortcomings with this study. Goodwin and Darley (2008) repeat the same problem in a second study, where they ask participants:
According to you, can there be a correct answer as to whether this statement is true? (p. 1351)
Once again, this is not a valid measure of realism. It is at best only capable of measuring cognitivism. To put it simply: Goodwin and Darley did not appear to understand the relevant conceptual distinctions, and designed studies that can’t plausibly tell us whether people are realists or antirealists.
Now let’s return to their other measure, the classic disagreement paradigm. Again, participants were asked to provide their response to a person disagreeing with them about a moral issue, and were asked to choose from one of the following options:
(1) The other person is surely mistaken.
(2) It is possible that neither you nor the other person is mistaken.
(3) It could be that you are mistaken, and the other person is correct.
(4) Other.
Once again, there are extremely serious problems with this. They interpret thinking that one person is correct and the other is incorrect as an indication of realism (R1 below), while thinking it’s possible neither person is mistaken as an indication of subjectivism (R2 below). Why is this a problem?
The main problem with Goodwin and Darley’s and Wainryb et al.’s methodology is again their inadequate metaethical assumptions. First, as they describe the moral disagreements at issue, R1 is not only entailed by realism, but also by various non-individualistic variants of subjectivism. Consider, for example, cultural relativism. According to this view, to judge a thing good means to judge that the members of the culture within which the judgement is made predominantly believe that the thing is good. Within one particular culture there can only be one predominant view about whether a thing is good. In order for cultural relativism not to entail that one party of a moral disagreement is right and the other is wrong, this disagreement must therefore take place between members of different cultures. However, neither group of researchers promoted such an interpretation. Quite the contrary! Goodwin and Darley described the disagreeing parties as subjects of their own study, which suggests that they are students of the very same university (see 2008, p. 1362). And Wainryb et al. even presented drawings which show the disagreeing parties standing face to face to each other (2004, p. 692).
In other words, the main response they used to indicate support for realism is entirely consistent with a nonrealist view! And to add insult to injury, their instructions actually encourage miscategorizing antirealist relativists as realists since participants are given information that would lead them to believe the people they disagree with are from the same culture; if they’re a cultural relativist, then they ought to judge that at least one of these people must be mistaken. In other words, the “realist” response is the appropriate response for cultural relativists. Could this have inflated “realist” response rates? Yes. Subsequent data from Pölzler and Wright (2020) suggests cultural relativism was the most common choice overall across multiple measures in their samples:
In reality, this is merely the tip of the iceberg. There are a legion of problems with their measures and instructions. The use of terms like “surely” is inappropriate because they introduce epistemic notions into the measures, which may lead to conflating what is intended to be a metaphysical question about the nature of moral truth with an epistemic question about what we’re in a position to know or are justified in believing. Likewise, language like “possible” and “could” can lead to unintended interpretations and introduce complicated modal notions.
This isn’t even getting into the titanic list of other problems with the study: the risk of normative/first-order conflations, epistemic conflations, the lack of appropriate response options for existing metaethical positions (e.g., noncognitivism), the failure to distinguish different forms of relativism/subjectivism, stimulus sampling problems, and the risk of unintended interpretations, including conflating realism and antirealism with absolutism, universalism, and descriptive claims. This list of issues isn’t even exhaustive nor is it speculative. Subsequent studies show that when you provide alternative metaethical options, they are frequently chosen, e.g., see Davis (2021) and Beebe (2015).
Likewise, my dissertation research shows that many of the unintended interpretations are frequently offered when participants are asked to interpret stimuli or explain their decisions. See Table 2.1 on page 43 of my dissertation where I present a three-page list of methodological problems with the disagreement paradigm. It is, if anything, remarkable that so many problems could apply to a single method; there’s a veritable cornucopia of problems, and if anything the disagreement paradigm would better serve as an instructional tool for how a measurement tool can catastrophically fail than a tool anyone should actually use. And, for what it’s worth, there is both direct and circumstantial evidence for the genuine relevance of many of these problems; they’re not merely speculative. In other instances, the problems are theoretical and relate to the operationalization of the construct, and are thus not the sort of problem to be assessed empirically in the first place.
It’s also worth noting another misleading way in which Goodwin and Darley interpreted their findings. Recall that they report that moral claims were about as realist as scientific claims, and were far more realist than matters of taste or social convention. There are two very serious problems with this. First, this ignores intradomain variation, i.e., the extent to which participants tended to take a realist or antirealist stance towards individual items within a given category (moral, scientific, etc.). As they themselves report:
Considering the ethical statements, it is noticeable that the assignment of truth to ethical statements varies considerably with the content of the statement. Participants generally agreed (on a six-point scale) with the goodness of anonymous donations (5.42), the badness of opening gunfire on a crowd (5.79), or of robbing a bank (5.77), and the wrongness of conscious racial discrimination (5.86) or of cheating on a lifeguard exam (5.72). But they varied considerably in how likely they were to regard these statements as true: 36%, 68%, 61%, 54%, and 58%, respectively. Perhaps more strikingly, although participants generally agreed (albeit not as strongly) with the permissibility of abortion (4.12), assisted death (4.36), and stem cell research (4.58) in the way we described them, they were highly reluctant to assign truth to statements expressing this agreement: 2%, 8%, and 2%, respectively. In other words, meta-ethical judgments about the truth of ethical claims appear to be highly sensitive to the content of the claims in question (i.e., robbery vs. abortion), and not merely to whether the claims are generally agreeable. Only 13 out of 50 participants applied the same category (truth vs. opinion) to the eight ethical statements they rated in the first part of the experiment. The remaining 37 varied their assignment of truth/falsity versus opinion in some way. (p. 1346)
In other words, they don’t find that scientific issues are generally regarded in realist terms, matters of taste and social conventions in antirealist terms, and moral issues in roughly realist terms. Rather, they found that people’s judgments were all over the place regarding moral issues, with massive inversions in the proportion taking a realist or antirealist stance towards different issues, or even the same issue. This should be a red flag. If you use two or three measures ostensibly intended to measure the same thing, and get wildly different measures for the same thing, this is an indication that one or more of your measures may not be valid. It’s at least an indication that they’re not measuring the same thing because they lack convergent validity.
However, there is a more serious issue, indicated in bold above:
In other words, meta-ethical judgments about the truth of ethical claims appear to be highly sensitive to the content of the claims in question (i.e., robbery vs. abortion), and not merely to whether the claims are generally agreeable.
Bingo! Most readers will be familiar with the importance of randomizing participants in a study. However, stimuli used to represent a category or domain are typically drawn nonrandomly from that category, yet this nonrandomness isn’t factored into their analyses. When researchers analyze data, they typically factor the randomness of participants into their analyses, but rarely consider or factor in that the same concerns apply to stimuli. Instead, they tend to employ nonrandom stimuli unsystematically drawn from the population of stimuli, then don’t bother to model this in their analyses. Is this a problem? Yes. Yes it is. See Judd et al. (2012). From their abstract:
Throughout social and cognitive psychology, participants are routinely asked to respond in some way to experimental stimuli that are thought to represent categories of theoretical interest. For instance, in measures of implicit attitudes, participants are primed with pictures of specific African American and White stimulus persons sampled in some way from possible stimuli that might have been used. Yet seldom is the sampling of stimuli taken into account in the analysis of the resulting data, in spite of numerous warnings about the perils of ignoring stimulus variation (Clark, 1973; Kenny, 1985; Wells & Windschitl, 1999). Part of this failure to attend to stimulus variation is due to the demands imposed by traditional analysis of variance procedures for the analysis of data when both participants and stimuli are treated as random factors.
Did Goodwin and Darley draw a random sample of moral issues to represent the moral domain?
No. They did not. They simply created ad hoc items for the study, with no systematic effort to evaluate the breadth or general representativeness of those items. Given this, it makes no sense to treat these ten items as representative of the moral domain. Why? For the same reason it wouldn’t be appropriate to use a sample of ten people you nonrandomly chose to represent everyone else. The whole point of randomization is to ensure representativeness. Simply put: if you swapped in ten different items, you might get completely different results.
Is this merely a speculative concern? No; it is not. On the contrary, it turns out that it’s incredibly easy to generate items that wildly shift the total proportion of people giving “realist” responses, not just to moral issues, but even to scientific issues. It would be easy, for instance, to intentionally generate a set of ten scientific items and ten moral items that actually gave the appearance of people being more realist about morality than science. Beebe (2015) demonstrated how this could be done in a brilliant study where he showed that irrelevant features of scientific claims led to dramatically reduced rates at which participants chose realist responses. All he had to do was present people with controversial scientific issues. Consider the realist rates for these scientific/factual issues.
Frequent exercise usually helps people to lose weight.
Global warming is due primarily to human activity (for example, the burning of fossil fuels).
Julius Caesar did not drink wine on his 21st birthday.
There is an even number of stars in the universe.
Humans evolved from more primitive primate species.
Mars is the smallest planet in the solar system.
The earth is only 6,000 years old.
New York City is further north than Los Angeles.
See the dramatic difference? Most people appear to be “scientific antirealists” when it comes to claims about exercise or climate change, and “realists” about the age of the earth or the location of NYC. Now, is it plausible these people are all “pluralists” who are scientific realists about uncontroversial scientific issues, but are antirealists about the truth of controversial issues?
No. It’s more plausible they’re not interpreting these questions to be about realism or antirealism at all. In other words, the most plausible interpretation of the disagreement paradigm is that it isn’t measuring what researchers think it’s measuring. I’m not merely suggesting this is plausible; I’m suggesting it’s the best interpretation. The alternative requires concluding some very strange things about the way ordinary people think. For instance, about half of the participants across every country tested gave a “relativist” response as to whether a historical event took place (e.g., Caesar drinking wine on his 21st birthday). Is it plausible that half of the world’s population are relativists about historical events, such that they think that whether a historical event did or didn’t happen depends on your personal preferences? No. This is ridiculous, and serves as a sanity check to anyone paying attention to the results of these studies.
Incidentally, Beebe’s findings also reveal the same massive variation in realist/antirealist responses to moral issues:
As you can see, if you happened to pick items more like donating to charity or euthanasia, your participants would appear to lean heavily towards antirealism, while if you chose issues like racism, they’d appear to favor realism. Since none of these items are representative of the moral domain as a whole (or at least, there’s no evidence they are, and they weren’t designed to be, and it’s very unlikely they’d be representative by chance), researchers who draw general inferences about the domain from which these items were drawn are committing an inferential mistake, the stimuli-as-fixed-effect fallacy:
The basic problem is that standard statistical analyses in psychology treat participants (subjects) as a random factor, but stimuli as a fixed factor. Thus our statistics assume that the goal of inference is to say something about some population that those participants are representative of (rather than just the particular people in our study). By treating stimuli as fixed it is assumed that we’ve exhaustively sampled the population of interest in our study. This limits statistical generalization to those particular stimuli. This is an unattractive property for psycholinguists (because they tend to be interested in, say, all concrete nouns rather than the 30 nouns used in the study). The same issue may apply to lots of other types of stimuli (faces, people, voices, pictures, logic problems and so forth). Baguley (2012)
In other words, researchers cannot make justified inferences about the moral domain based on nonrandomly sampled stimuli. Since the items used by Goodwin and Darley weren’t randomly sampled, they cannot be used to make inferences about the moral domain as a whole. They simply aren’t in a position to claim that moral claims are treated almost as objectively as scientific claims, even if their data showed this, which it doesn’t. At best, they could only say that the haphazard set of moral issues they happened to choose averaged out in a way intermediate between how the scientific and taste issues they also haphazardly chose happened to average out. They can’t justifiably generalize from their stimuli any more than they could justifiably generalize from a nonrandom sample of participants to broader populations. As such, from the very ground up, their study design is simply inappropriate for making generalizations about whether people treat moral issues in general in a way that leans towards moral realism.
There are ways to mitigate this problem, but they are fairly difficult. In Moss et al. (2025), my colleagues and I devised a much larger set of moral issues that vary across a range of dimensions relevant to the degree to which these items represent the moral domain as a whole:
While prior studies have also employed a variety of different moral issues as items when examining individuals’ metaethical judgment (Goodwin & Darley, 2008, 2012; Pölzler et al., 2022; Pölzler & Wright, 2020), for the most part, these studies have not systematically developed selections of items which vary across a set of different dimensions. For this reason, we created a database of 44 items with the a priori intent, based on our theoretical judgment, to vary along the dimensions of severity (low, medium, high), perceived objectivity (low, medium, high), and moral foundation (harm, fairness, authority, loyalty, purity). To examine how these items (along with an additional item that was removed in the study) varied in different dimensions (such as perceived objectivity, agreement and perceived consensus) we conducted a pretest with 300 participants. The pretest showed that the items vary substantially in these dimensions.
While this approach isn’t a perfect solution to the problem, it at least serves to mitigate problems with stimulus sampling. This is one small part of what it looks like to address methodological shortcomings with earlier research, and it should illustrate that I don’t take these methodological problems to be idle or merely abstract; circumventing them requires real effort.
Again, I must emphasize: Rian mentions none of this. He just presents the most superficial summary of the findings of the studies he mentions possible, with no appreciation for the actual content of these studies. He doesn’t carefully engage with their methods, their results, how those results were interpreted, or any other factors relevant to assessing the actual merits of the studies.
And again, this is before we even get to the issues David Moss and I focus on in the 2020 paper. In other words, before we even get to what I consider the most serious issues to be, this study already has so many problems there is very little reason to take the results very seriously. I don’t say this because I hate the study or think it’s worthless. I really appreciate this research and think they did a great job. But doing a great job can still result in a failure. We learn from those failures, make changes, and move forward. That’s how science works. Careful knowledge of the actual content of the science requires considerable effort and deep familiarity. Rian shows zero indication of serious engagement with the literature.
Now, let’s address the main issue Moss and I focus on: unintended interpretations. Rian says:
including, ofc, your own work with Moss - which I personally find pretty weak, since showing that standard probes are open to non-metaethical interpretations does not by itself establish that they systematically fail to track metaethical attitudes and at most it raises a validity concern […]
Our article doesn’t merely show that alternative interpretations were possible; we provide direct evidence of their frequency. Yes, our findings do show that their study systematically failed to track metaethical attitudes for a substantial proportion of participants. We don’t even need to do this though, because the studies in question were so poorly operationalized and so severely misinterpreted that if anything the onus would be on the authors to provide any good reasons to think their studies did track the intended metaethical attitudes in the first place. But how that’d even be possible should be a complete mystery, given that most of their measures aren’t even face valid (i.e., they don’t even appear to measure what they’re intended to).
I reanalyzed the Goodwin and Darley data again for my dissertation, so I’ll talk about that instead of the earlier paper. Here’s what I did. Goodwin and Darley happened to gather open response data from participants where they asked those participants to explain what they thought the source of the disagreement between themselves and the other person might be. They reported that virtually everyone interpreted the source of disagreement as intended, which would raise few validity concerns. I was skeptical of this, and evaluated the data myself. I found a clear intended interpretation rate of only 25%, and a clear unintended interpretation rate of 44% for their first study, and 36.2% intended and 52% unintended for their second study. Note that these proportions are asymmetric: interpreting the source of disagreement as intended is a necessary condition for the measure to be valid, but it isn’t sufficient. Conversely, if they didn’t interpret the source of the disagreement as intended, this would be sufficient to entail that their answers were invalid.
Because of this, a clear unintended interpretation would invalidate the results for a given participant, but a clear intended interpretation wouldn’t, since it reflects only one way in which they may have interpreted the overall stimuli as intended. Thus, 25% and 36% are at best upper bounds on how many people were likely to have interpreted the stimuli as intended (this is setting aside that some significant proportion of unclear responses might have interpreted the stimuli as intended, but this may be compensated by the sheer number of other problems that drive plausible unintended interpretations). Either way, the rate of clear unintended interpretations is high enough to be catastrophic to the validity of these studies (which, again, probably aren’t valid for other reasons, anyway).
I also replicated these findings in a follow-up study, Study 1C, in section S4.4.1 on page S4-590 of my dissertation. Finally, I conducted thematic analysis on the various responses people gave, which you can see beginning on S4.7.3, S4-670 of the dissertation. This involves devising distinct qualitative categories and counting up the number of responses that fit into one or more of these categories. Here you can see the way I classified the various responses, many of which clearly involve unintended interpretations. These unintended interpretations are so frequent that yes, we did provide pretty good evidence that their measures systematically fail to track metaethical attitudes, and again, given that they misoperationalized their measures from the outset, there are already sufficient theoretical grounds for doubting that they did so in the first place.
That, at the very least, is a more plausible conclusion than that their measures somehow do track the intended metaethical attitudes in spite of their poor face validity, stimulus sampling problems, misleading instructions, use of a forced choice paradigm that introduces confounds in the response options, and so on. Theoretical considerations are often sufficient on their own to cast serious doubt on a study, and this is one such case, entirely independent of my empirical findings. Collectively, these considerations are more than enough to firmly shift the burden squarely on proponents of the original study to demonstrate that their measures somehow did track metaethical attitudes as intended, despite so many people saying things in response that suggested they didn’t interpret the stimuli as intended.
Now, just because a person attributes the source of disagreement to something other than a fundamental moral disagreement, this doesn’t necessarily mean their results are invalid. This is because these open response questions rely on the participant giving a post hoc explanation of their reasoning and interpretation. It’s always possible that they did interpret the stimuli as intended at the time of their judgment, but their post hoc response doesn’t reflect this. I think this can and does occur, and probably quite often. So the percentages above might underestimate the amount of intended interpretations. But if a person answers in a way that indicates they didn’t interpret the question as intended, it isn’t reasonable to shrug and assume the measure was probably valid in spite of this; at the very least, the very high rate of unintended interpretations shifts the burden onto those who claim the measures are valid to provide evidence of their validity.
In other words, are David’s and my findings so strong we’ve established, definitively, that Goodwin and Darley’s measures are not valid? No. Nor have I provided overwhelming or even very strong evidence in other cases. My findings are fairly tentative. But seeing as most of the studies in question offer virtually no evidence or very little evidence in support of the validity of their measures, and given that there are good theoretical grounds for doubting the validity of these measures, the collective consequence of these theoretical and empirical issues is at least sufficient to shift the burden onto proponents of the validity of these measures. Could they be valid in spite of the concerns I and others have raised? Sure. Has their validity been substantiated? No.
The validity concerns I raise are not minor. They are extensive. The entire second chapter of my dissertation addresses more than a dozen other problems with the disagreement paradigm. The issues are genuinely that extensive.
4.0 Nichols & Folds-Bennett (2003)
An entire section of my dissertation is specifically dedicated to the methodological problems with this particular study. See the supplement to chapter 3, S3.4, page S3-396. This study also doesn’t provide good evidence that people are realists. I don’t want to have to fully repeat myself, so go check that section for a comprehensive discussion of the problems with this study. Here, I’ll note a few points:
First, the study was only conducted with children. Rates of realist and antirealist responses change over the lifespan and longstanding research in developmental moral psychology shows that children often give quite different responses than adults. Researchers cannot meaningfully generalize from studies in children to studies in adults, so even if this study found that most children were moral realists, this wouldn’t provide good evidence that moral realism was common among adults.
This isn’t really a methodological problem so much as a limitation worth noting. The issues with the study come from its actual design. The goal of the study is to determine whether children consider moral concepts like “good” and “bad” to be distinctively non-response-dependent, in contrast to taste preferences, which presumably are response-dependent. What does this mean? As they put it:
the basic idea is that a property is response-dependent just in case that property is constituted by the responses it elicits in a population; so the same object or event might have different response-dependent properties in different populations (p. B25).
“Icky” serves as an example. Whether something is icky or not depends on the responses of different individuals, and can therefore vary. Conversely, one might suppose that a property like “round” wouldn’t depend on responses in the same way. The question then is whether children treat moral concepts as response-dependent or not. If they do, this would imply they are antirealists, while if they don’t, this would imply that they’re realists. So what did they do to test this claim?
Children were asked whether a particular food like grapes was “yummy,” then given a scripted response intended to determine whether they treated taste preferences as response-dependent:
You know, I think grapes are yummy too. Some people don’t like grapes. They don’t think grapes are yummy. Would you say that grapes are yummy for some people or that they’re yummy for real? (p. B27)
Children are thus given a dichotomous choice between the following response options:
Yummy for some people
Yummy for real
There are a few problems with this setup. First, these options are not mutually exclusive. The fact that something is yummy for some people does not clearly entail that it’s not yummy for real. Consider, for instance, the fact that peanut allergies are dangerous for some people. Does this mean they’re not dangerous “for real”? Of course not. From the very outset, then, the response options aren’t designed very well. They already imply that for something to be yummy for only some people rather than everyone somehow means that it isn’t real. This doesn’t make much sense.
It also implies that response-dependent properties are somehow not real, and that only response-independent properties are. But this isn’t a requirement for endorsing one of these views, so the language used isn’t appropriate for its intended purpose. Realists are not, after all, required or even necessarily inclined to think that response-dependent or nonrealist properties aren’t real. If anything, such a characterization would reflect a bias against or misunderstanding of antirealist views; presumably response options shouldn’t be designed with the presumption that everyone who holds such a view is biased or confused about opposing views.
Furthermore, the only form of response-dependence tested here is personal response-dependence. While this makes sense for matters of taste, e.g., the notion of “icky,” if parallel structure was used for moral items this would mean that the only contrast was between a form of direct, individual relativism and the putative contrast of realism. But this is not an appropriate way to distinguish realism from antirealist options: it essentially involves giving one an option between what the study operationalizes as realism and what effectively serves as a narrow and specific form of relativism. Yet there are other forms of stance-dependent cognitivist accounts: cultural relativism, constructivism, and so on, according to which moral truths are stance-dependent but their truth doesn’t directly depend on our personal response to them. Given this, the study has misoperationalized the response options in such a way that isn’t genuinely exhaustive of the range of potential stances one might take towards them.
And that brings us to the more serious problem: this measure is only valid if children interpreted the question and response options in a way that if they did hold realist views, they would reliably choose “for real” while they wouldn’t if they didn’t. But why should we think this is a valid way to distinguish realists from antirealists? More generally, what does it even mean to say that grapes are yummy for real? Why should we suppose children would specifically interpret this to mean that whether grapes are yummy doesn’t depend on our reaction to their taste? Why would that be the only sense in which something could be yummy “for real”? Consider your own taste preferences. Do you like or dislike pineapple on pizza? Whatever your preference, if you think that whether pineapple on pizza is yummy or not depends in some way on your personal preferences, and that whether it’s yummy can vary from one person to another (i.e., it’s response-dependent), does this mean that you don’t think pineapple on pizza is good or bad “for real”? Are matters of personal taste somehow not real?
Furthermore, note that this question presents a forced choice. Your only options are to choose that it’s yummy for some people or for real. But who would disagree that grapes are yummy for some people? Of course they are! This is uncontroversially true! So why does the question present a forced choice between an unobjectionable claim and another, more obscure claim that isn’t even inconsistent with it? There simply isn’t a proper dichotomy or set of exclusive choices. But the pairing of “for some people” with another, more obscure option may prompt the presumptive contrast to be that they are either yummy only for some people or for everyone. If so, this would imply universality, and would instead be interpreted as a question about the proportion of people who like grapes. If so, children may have interpreted this as a question that effectively asked whether only a few people like grapes or if most or everyone does. Finally, “for real” can sometimes be used as a form of emphasis. If so, it could be interpreted as a remark about how yummy grapes are.
Either way, children may have judged that preferences about grapes are more variable than our stances on moral issues, or judged that taste preferences aren’t as serious and thus worth emphasizing, both of which could explain the asymmetry between children choosing the “for some” response for matters of taste to a greater extent than they did for moral issues.
In addition, it’s worth noting that this study also used a small pool of ad hoc stimuli and thus may likewise suffer from stimulus sampling issues. Nichols’s (2002) own work has shown that when the range of social/conventional violations is expanded to include violations previous studies hadn’t employed, the clustering of characteristics typically associated with the moral/conventional distinction came apart. This was an excellent illustration on his part of the problem of relying on a narrow and repeated subset of stimuli. In other words, one of the authors of this study has already conducted studies that reveal stimulus sampling problems in a related line of research. Why think such problems simply wouldn’t arise here?
Now, are these points speculative? Sure. And these possibilities aren’t exhaustive of the range of plausible alternative interpretations that could have prompted children to respond to the stimuli in ways unrelated to them taking any position on response-dependence/independence. But notice an asymmetry here. Is this how we judge whether a measure is valid: assume the measure is valid unless someone demonstrates that it isn’t? No. Measurement is not innocent-until-proven-guilty. If one can raise legitimate theoretical reasons to be skeptical of the validity of a measure, these can be sufficient to shift the burden onto those who employ the measure to demonstrate that it’s actually valid.
And it’s not like there are no tools or methods available to assess validity. You can assess validity by assessing whether different measures of the same construct yield similar results, you can consult expert judges, you can employ comprehension checks in your studies, you can use qualitative methods to assess how stimuli were interpreted, you can use cognitive interviewing to assess how participants are thinking through stimuli in real time, and so on. As far as I can tell, none of these methods were mentioned in the paper and thus it’s unlikely any active efforts were made to validate the measures that were used. I also don’t know of any attempts to replicate these results or corroborate the validity of these methods elsewhere. This is not an idle issue. Lack of evidence of validation is a serious and widespread problem in psychological research. See e.g., Flake et al. (2017):
The verity of results about a psychological construct hinges on the validity of its measurement, making construct validation a fundamental methodology to the scientific process. We reviewed a representative sample of articles published in the Journal of Personality and Social Psychology for construct validity evidence. We report that latent variable measurement, in which responses to items are used to represent a construct, is pervasive in social and personality research. However, the field does not appear to be engaged in best practices for ongoing construct validation. We found that validity evidence of existing and author-developed scales was lacking, with coefficient α often being the only psychometric evidence reported. We provide a discussion of why the construct validation framework is important for social and personality researchers and recommendations for improving practice.
Now let’s set aside all of the preceding issues. There’s still one very serious problem with this paper: the extremely low sample sizes for the studies, just n = 19 and n = 13 for their two studies. I cannot stress enough that these are extremely low sample sizes. These sample sizes would only be suitable for reliably detecting extremely large effects, and most effects in the social sciences are fairly modest. To my knowledge, this study has not been replicated, nor has their response-dependence paradigm been employed elsewhere. It’s likely that if there is a real effect, it is fairly small. Given this, both the results of this study and the methods used should be interpreted with considerable caution. This pair of studies would certainly not be adequate to justify any confident conclusions that most children are moral realists. While the study’s design is innovative and the method they take should be refined and replicated, this study hardly contributes to any robust body of empirical data showing that an appreciable number of people find moral realism “extremely intuitive.” In short: even if we set aside every methodological issue with the study and assumed the results were entirely without issue, these findings would provide extremely marginal evidence that realism is “extremely intuitive” even within the tested population, much less “across populations.”
Once again, it’s remarkable Rian would cite this study of all studies as support for the notion that we have good evidence of folk realism. This is, like the Goodwin and Darley study, one of the weaker studies, if for no other reason than its extremely low sample size, not to mention the use of a one-off paradigm in a sample of children that wasn’t replicated with a larger sample or across cultures. There are other, better choices for studies that serve as evidence for widespread realist intuitions. For instance, if you wanted a better developmental study, Reinecke and Heiphetz Solomon (2023) would be a better choice. While its measures are indirect, it at least features a much larger sample size (n = 129). However, there are also reasons to be skeptical that this study provides strong evidence that children are realists. See here.
Likewise, if you want an adult study with a broader range of measures and that has fewer methodological problems than Goodwin and Darley (2008), you could opt for e.g., Zijlstra (2023). This study also has significant shortcomings, but if I were trying to present evidence that realism was extremely intuitive I’d probably opt for something like this, and maybe Wright, Grandjean, and McWhite (2013) since it serves to replicate Goodwin and Darley’s findings under an additional constraint, which makes it at least slightly stronger. The point here is that there are better papers one could use as an example, but Rian didn’t use them. This suggests Rian really just isn’t very familiar with the academic literature. The articles he presents appear to have been chosen in an unprincipled way, and, ironically, are some of the weaker studies in support of his position.
5.0 Beebe et al. (2015)
Beebe et al. (2015) is the best of the three studies Rian cites, but it does little to advance the case that we have good evidence that moral realism is a highly intuitive position across cultures. Unlike the other studies, it actually presents cross-cultural data. However, this study employs the disagreement paradigm, and inherits many of the problems with Goodwin and Darley’s studies (though fortunately not the worst of them).
The main problem is that it still employs the disagreement paradigm. As I argue in my dissertation, this method has well over a dozen problems. Its invalidity is overdetermined by the sheer number of flaws with this method, foremost among them low intended interpretation rates.
It still has the stimulus sampling problem Goodwin and Darley’s studies suffer from; there’s no evidence or good reason to believe the set of moral items used represent the moral domain, and as such generalizing to other moral issues or the moral domain as a whole is inappropriate.
A couple other points worth noting about this study. First, the results show that realism rates are not very impressive across the moral items that they did present:
Averaging across the non-US samples, realist rates range from 33% to 67%, two items fall below the mean, and no items even approach unanimity. Rian never made it clear what it meant to say that realism is extremely intuitive across populations; if this just means it’s common then this would be evidence for that, but it wouldn’t support stronger claims about it approaching unanimity or being a strong majority view.
However, one reason it’s a bit ironic to choose this study is because their findings highlighted one of the most serious problems with the disagreement paradigm. Note in the graph above how low realist rates are for matters of scientific fact. The 21st birthday is a reference to whether Julius Caesar (or Confucius) ate soup (or drank wine; it varies across studies) on their 21st birthday. Consider this question set up:
If someone disagrees with you about whether Julius Caesar ate soup on his 21st birthday, is it possible for both of you to be correct or must one of you be mistaken?
□ It is possible for both of you to be correct
□ At least one of you must be mistaken
If this question was interpreted as intended, selecting the first option would entail relativism about historical events. This would mean that it would be possible for a historical event to have both occurred and not occurred depending on the personal attitudes of the individual considering the matter. Across every culture, Beebe et al. found that about half of participants chose the “relativist” response to this question.
Is it even remotely plausible that half of the world’s population (or at least these cultures, assuming they’re not representative) are relativists about historical events?
No. It isn’t. The most plausible conclusion is that people did not interpret this question as intended. This wasn’t a one-off anomaly for this particular item. Realism rates were alarmingly low for the other factual issues they employed. These items ought to serve effectively as a comprehension check: not selecting the realist response option is better evidence that participants didn’t interpret stimuli as intended than it is evidence that the participant is a relativist about the issues in question. Take, for instance, the claim “Frequent exercise usually helps people to lose weight.” Realist rates for this were 38%, 43%, and 40% in the non-US samples. Is it remotely plausible that most of the participants think there is no stance-independent fact of the matter about this question? No, it isn’t.
The disagreement paradigm is not a valid measurement tool. Again, this is the best Rian has? Two weak studies and one study that, if anything, actually provides evidence against the main method used in experimental metaethics.
6.0 Sousa et al. (2021)
Now this brings us to Rian’s flagship study. As we will see, this study does nothing to help Rian and serves as evidence only of the fact that Rian doesn’t know what he’s talking about. Rian comments:
[…] including, ofc, your own work with Moss - which I personally find pretty weak, since showing that standard probes are open to non-metaethical interpretations does not by itself establish that they systematically fail to track metaethical attitudes and at most it raises a validity concern, and revised later work has explicitly tried to control for alternative interpretations and still found high rates of objectivist responding, especially for clear cases of harm and injustice, but anyway). More recent work such as Sousa et al has in turn argued that some apparently relativist results may partly reflect those very measurement problems.
This is a reference to Sousa et al. (2021). Fortunately, this article is freely available so you should be able to see the whole thing here. As you can see, Rian’s remarks imply that this article represents “revised later work” that addresses the methodological issues raised by Moss and me. Which work of mine is Rian referring to? He’s not clear, but his remarks imply it includes my dissertation, which would be a bit puzzling, but note what he says in response to Nathan, who points out that:
Lance’s entire dissertation is about how these studies don’t actually support folk moral realism
Rian responds:
You didn’t bother to read the whole text, right?
Rian is clearly suggesting Sousa et al. somehow addresses the concerns raised in my dissertation. It does not. It doesn’t even attempt to do so. Which is not surprising: it couldn’t have, given that my dissertation came out after this article! How could they have possibly addressed the specific concerns I raised before I raised them? Perhaps they noticed the same problems before my dissertation was published, then sought to address them? No. They didn’t. There is no indication of this in the article at all. They didn’t cite Bush and Moss (2020), nor did they show any particular interest in addressing the specific issues we raise or that I raise in later work.
This is because the methodological issue they address was one raised by a specific paper, Sarkissian et al. (2011). The goal of Sousa et al. is only to address a methodological issue associated with that particular study. The issue in question isn’t one I focus on in Bush and Moss (2020) nor is it an issue I focus on in my dissertation. Sure, they try to advance a case that many people are realists about a certain subset of issues, but they do so relying on—and failing to correct for—the problems with the same methods that I focus on in my work. The notion that this article somehow addresses my concerns makes absolutely no sense, and is a clear indicator that Rian may not even have a superficial understanding of the literature; he may be actively confused about it. But don’t take my word for it. Let’s have a look at the article.
The primary goal of the article is to establish that nonphilosophers at least tend to be realists about harmful actions that involve injustice. They attempt to establish this by developing on and reacting to a study by Sarkissian et al. (2011), so let’s talk about that study. Sarkissian et al. (2011) used a version of the disagreement paradigm. However, they also varied the degree of cultural difference between the participant and the hypothetical person who disagreed with them. They employed three conditions in which the person who disagrees with them is:
From the same culture
From an Amazonian warrior culture
From an alien species called Pentars whose only interest is in converting all matter into pentagons
What they found is that realism levels were highest in the same culture condition, intermediate in the other culture condition, and lowest in the other species condition. Their proposed explanation for this is that people don’t have a fixed commitment to moral realism or antirealism. Instead, people’s commitment varies in accord with the degree to which radically different moral perspectives are salient (i.e., present in awareness).
According to Sousa et al. (2021), this study has some methodological limitations:
In particular, participants may have assumed that the exotic or alien party misunderstood the harmful action, and this assumption, rather than a genuinely non-objectivist stance, may have contributed to the increase in non-objectivist responses.
This is a reasonable proposal. In order for the disagreement paradigm to serve as a valid measure, both of the people who disagree (the participant and whoever else) must understand the moral issue in the same way. If participants assume the other person misunderstood the scenario, then this may account for the participant’s response rather than the participant judging that the other person’s moral position is also correct. You can see that their entire emphasis is on addressing the question of whether the person from the other culture or the alien may have misunderstood the scenario, and their primary objective is to minimize the impact of this particular possibility on study results:
Study 1 replicated Sarkissian et al.'s results with additional follow-up measures probing participants' assumptions about how the exotic or alien party understood the harmful action, which supported our suspicion that their results are inconclusive and therefore do not constitute reliable evidence against the deflationary hypothesis. Studies 2 and 3 modified Sarkissian et al.'s design to provide a clear-cut and reliable test of the deflationary hypothesis. In Study 2, we addressed potential issues with their design, including those concerning participants' assumptions about how the exotic or alien party understood the harmful action. In Study 3, we manipulated the alien party's capacity to understand the harmful action. With these changes to the design, high rates of objectivism emerged, consistent with the deflationary hypothesis.
Before we have a look at what they did in their studies, let’s have a look at a few remarks made in their literature review:
Many results using this methodology (and ones like it) are indeed consistent with our hypothesis that people are objectivists concerning unjust harmful actions—e.g., cheating, stealing (for a review, see Goodwin, 2018). The great majority of adult participants accept that only one of two incompatible beliefs is correct with regards to such actions (see, e.g., Nichols, 2004; Goodwin and Darley, 2008, 2012). Furthermore, developmental work with children aged 5–13 years has shown that children are even more objectivist than adults in this respect (see, e.g., Wainryb et al., 1998, 2004). But while most current studies using the incompatible-beliefs paradigm provide clearer evidence about objectivism, they fall short of providing evidence on whether the objectivism at stake is one with a universalist scope, since the action evaluated is normally depicted in the context or society of the participant, rather than a context or society different from the participant’s own. As a result, these studies too fall short of providing complete evidence for the deflationary hypothesis.
Notice how the authors uncritically accept studies that employ the disagreement paradigm as generally valid and as establishing the ubiquity of moral realism among both children and adult populations. This would be a very strange set of remarks to make if the authors were highly sensitive to the suite of methodological issues raised by Pölzler, Moss, and myself among others. Notice how they present the Sarkissian study immediately following this:
Be that as it may, Sarkissian et al. (2011) provided evidence that seems to undermine the deflationary hypothesis. They argued that existing findings using the incompatible-beliefs paradigm are somewhat misleading because the incompatible beliefs are attributed to appraisers from the same cultural background or psychological profile.
Their focus, then, is only on addressing this one challenge to existing results. Even if they completely overcame this challenge, it would not entail that they addressed any of the challenges I’ve raised. Why? Because this isn’t even one of the challenges I raised, and even if it were, it’d only be one of them!
Simply put: this study has nothing to do with addressing the methodological problems I’ve raised. The authors begin with the assumption that the disagreement paradigm is a valid measure other than to the extent that it’s threatened by Sarkissian et al. (2011); they then seek to address the methodological problems raised by that article, and then conclude that because those concerns have been addressed, they’ve therefore established their conclusion that available evidence suggests people are moral realists at least about a subset of moral issues. The reasoning is all tightly contained within this structure, and in no way grapples with the concerns I raise.
Their objectives are, in fact, even more narrow than this. As they say:
We think Sarkissian et al. (2011) were right to identify limitations in the previous research, and that the overall strategy they used to remedy this was pertinent and original. Nonetheless, we do not think their results decisively resolve matters owing to potential issues concerning participants' interpretations of the disagreements. In their studies, the description of the stabbing action, reproduced above, does not completely specify all of the morally relevant aspects of the action (i.e., those aspects bearing on the perceived injustice of the act), which leaves open the possibility that the harmful action was understood as an instance of justifiable harm, rather than as involving injustice. This may not have affected participants' own understanding of the action, or their interpretation of the American appraisers' understanding of the action. But it may have led many participants to infer that the appraisers from a different culture or species had a very different understanding of the situation. For instance, participants may have inferred that, given their radically distinct background, the exotic and alien appraisers understood the harmful action as being one that was performed with the consent of the victim (e.g., as part of a cultural ritual), or because the victim was guilty of a crime, or for some other justifiable reason. In other words, the exotic and alien appraisers may be thought to have appraised fundamentally different harmful actions, actions that might be justified, even from the point of view of the participant.
Their focus was thus not only on specific conditions used in Sarkissian et al. (2011), which it’s worth noting are virtually never employed in any other studies, but on the specific moral issues employed alongside those conditions. It doesn’t get more narrow than this. Debunking problems posed by one study doesn’t, in any way, show that there aren’t any other problems. And again, this isn’t even one of the problems I focus on.
Yet for some reason, Rian seems to think this particular study somehow resolves the concerns I raise in my work. There is no way for Rian to think this unless he is comically confused or simply didn’t read the Sousa article at all.
This brings us to Rian’s final post, which you can find here. Since it’s very long, I’ll handle it in parts.
Here we go again! Lance “Burro” (as we call here in Brazil - in case anyone in the audience are reading this and don’t know what “burro” means, feel free to Google it) strikes again.
My wife and in-laws are Brazilian. Burro literally translates to “donkey” and means something like “idiot” or “stupid person.”
I’m going to skip over the implicit ridiculous appeal to authority at the beginning of your post.
There is nothing ridiculous about pointing out that I am an authority on this topic. There’s nothing ridiculous about drawing attention to one’s expertise on a topic. Appeals to authority only become a problem when one appeals to authorities in a way that serves to circumvent the evidential or argumentative foundation on which that authority rests, e.g., “you should accept this claim because it comes from authority and for that reason alone.” I didn’t do that, so it’s unclear what exactly Rian thinks is ridiculous. Rian may have a poor understanding of informal fallacies and mistakenly and reflexively react to any reference to authority or expertise as somehow ridiculous.
In any case, I am not making an implicit appeal to my authority; I’m making an explicit appeal. I specifically specialize in the psychological methods used to assess whether nonphilosophers are moral realists or antirealists. I obtained my PhD in psychology in 2023 and wrote my dissertation specifically on this topic. I’ve spent over a decade studying experimental metaethics and the methods used in the field. While there are a handful of researchers who have published in the field, very few specialize in it. I am one of the only people in the world who specializes in this topic. It is entirely possible that there isn’t a single person out of the ~8.3 billion people on the planet more knowledgeable than me about this particular body of literature.
So when I say I’m an expert on this literature, this is, if anything, an understatement. A reasonable case could be made that I am the expert on this research. Does that make me perfect, or omniscient, or infallible? No. But this isn’t a standard case in which a person has a degree in a topic and then points to the degree as evidence they’re an expert on the topic. Rather, this is a case where a person has dedicated a considerable portion of their adult life to specifically studying this particular topic. It’s a niche field, there aren’t that many people in it, many of the people who published in this area have moved on to other topics, and it’d be exceptionally difficult to find anyone else with more extensive knowledge of these studies than myself. I cannot think of a single plausible candidate. They might be out there, but they’re certainly keeping their head down. This isn’t a matter of being pompous or arrogant or thinking I’m superior to other people. I always do my best to defer to the facts and data, and I have a long history of opposition to gatekeeping, elitism, and credentialism. So what I am not doing is saying something like “I study this topic and have degrees, therefore I am correct about it.” Far from it.
What I am saying is that if anyone is looking at this conversation from the outside, it is far closer to a situation in which a random person on the internet is arguing with one of the world’s foremost specialists in a particular disease about the nature of that specific disease than it is to an argument between two peers or even an argument between a random person and a random doctor about a general matter of health. Should people defer to me merely because of my expertise? No. I don’t want anyone to do so, including Rian. If I did, I wouldn’t have bothered writing all of this. I’d just appeal to my credentials. But I didn’t do that. And I didn’t take people’s advice to ignore people like Rian, because I dislike the notion that questions should be settled by appeals to credentials or authority and dislike precisely the mindset where a person would deem others unworthy of their time or attention merely due to their lack of credentials. I don’t personally care very much about credentials at all. But I will note that, in this particular case, I have them. That’s not really what matters, though. What matters is that I can back them up.
Questions should be settled by an examination of the actual basis for our respective views: the arguments and evidence we’re capable of marshalling for our respective sides. So that’s what I am doing throughout the entirety of the rest of this post. But I do think it’s worth flagging that I am, in fact, an expert in this research, and that Rian almost certainly isn’t.
Let’s recall what your original claim was: that there is no (good) evidence whatsoever for the claim that something like moral realism is intuitive across populations and that anyone who says otherwise is simply asserting it gratuitously, out of thin air, in an unsubstantiated way.
That is not my original claim. Rian has added and embellished on the original remark. This is a recurring theme in Rian’s responses: he twists, distorts, and exaggerates other people’s claims. This is my original claim:
There is no good evidence that moral realism is “extremely intuitive across populations.” It’s downright bizarre for people to keep making this unsubstantiated empirical claim.
Notice how Rian tosses my use of the term “good” into parentheses, which serves to downplay the critical importance that qualification plays in my original remark: my position is that there’s no good evidence. It shouldn’t be in a parenthetical. I also don’t use the phrase “whatsoever.” This serves to subtly distort my claim, especially when appended to “no” with “good” shunted into a parenthetical. It makes it sound a bit closer to the notion that I’m claiming there’s no evidence at all. I also never said that anyone who says otherwise is doing so in a way that is gratuitous or “out of thin air.” Rian simply made these up. That’s a bit ironic, for someone who has made a point of accusing me of dishonesty. Is it dishonest to say that you’re going to present someone else’s claim, then embellish and distort that claim by adding elements to it the person didn’t say? Only if you’re doing so intentionally, as far as I’m concerned. And unlike Rian, I don’t take myself to have the telepathic power to tell if a person is lying.
Now, anyone who has taken Epistemology 101 knows that evidence can only be called “good” in the relevant sense insofar as it makes it rational to hold some belief X.
This isn’t Epistemology 101. It’s not even true! This is one way in which a person might use the term good, but it isn’t the sense in which I was using it. But a person can adopt a more modest conception of what it would mean for there to be good evidence. On this more modest view, good evidence is something more like evidence that convincingly and significantly raises the probability of a given claim, all else being equal. This is consistent with it still not being rational to hold the view in question. This is because there can be very good evidence for X, but such good evidence against X that it isn’t rational to believe X, anyway. Whether a belief is rational or not turns on consideration of all relevant evidence available to you, not just isolated pieces of evidence.
When I made the claim that there’s no good evidence, I wasn’t merely claiming the evidence wasn’t good enough to justify Rian’s claim, I was making the stronger claim that there isn’t even a single study that significantly raises the probability of the claim, all else being equal. Ironically, Rian seems convinced I’m lying about a claim that he thinks conflicts with the data, but the actual claim I’m making is, in fact, stronger than the claim Rian seems to think I’m making.
So Rian is wrong about what I meant and fails to acknowledge alternative interpretations of what I said, one of which is more accurate and makes an even stronger claim than Rian supposes. Rian continues:
For example, atheism/naturalism is, in my view, a false belief. That doesn't mean I say that there are no good justification/evidence for atheism/naturalism, given, for example, the evidential problem of evil. It's simply the case that, since long before Gettier, everyone should know that there's a distinction between being justified/rational in believing something (that's all substantiation/evidence/justification could be about) and that belief being true.
This latter remark is presumably intended to imply I don’t understand this distinction. But it is Rian who is confused. You could hold the following views about an opposing position:
There is no evidence for the view
There is some evidence for the view, but it isn’t very good and doesn’t justify belief in that view
There is some evidence for a view, and it is good, but it’s not enough to justify belief in that view
There is some evidence for a view, and it is good, and it’s enough to justify believing in that view (even though you yourself don’t believe it)
Nobody is obligated to think that if there’s evidence for a view they don’t hold that the evidence is good, nor do they have to think the evidence is good enough that belief in the opposing view is rational. They can think this, but they don’t have to. It almost seems as though Rian doesn’t appreciate that a person can endorse (2) or (3) above. Sure, maybe Rian thinks there’s some good evidence for naturalism/atheism, good enough to justify belief in these views. That’s fine.
But I don’t think there’s good enough evidence that it’s rational to think moral realism is “extremely intuitive across populations.” I don’t even think there’s any good evidence in the non-justificatory sense!
What all of your works shows is that there are objections to these empirical studies - especially concerning the metaethical interpretation inferred from them. Fine. That happens with every conceivable piece of empirical research. It would be strange if it couldn’t. What none of your work shows: that there is no good reason whatsoever to believe that there is empirical evidence substantiating the claim that moral realism is intuitive across populations and that everyone who's saying this is mere affirming in a unsubstantied manner.
It apparently hasn’t occurred to Rian that I don’t agree with his assessment of what my research shows. No, my data doesn’t merely show that there are objections to the empirical studies. My arguments and data show that there are numerous compelling objections to the validity of existing methods in experimental metaethics and that, as a result, we should conclude that none of the measures used in these studies are valid, which in turn indicates that claims that moral realism is widespread among nonphilosophers are unsubstantiated. My work, in other words, does show that there’s no good reason to believe moral realism is intuitive across populations and it does show that claims to the contrary are unsubstantiated. The scope and severity of the issues I raise don’t generally apply to other psychological research, so Rian’s claim that my objections are applicable to “every conceivable piece of empirical research” is simply untrue. The sorts of validity issues I raise do apply to some other paradigms, but many are distinctive to metaethics or to a subset of other, similarly flawed methods, and Rian doesn’t present any good reason to think they do, anyway.
In fact, I’ll go further and say that Rian’s interpretation of my data is itself unsubstantiated. Note that Rian simply asserts these claims about my data and arguments. He doesn’t quote a single piece of my work that I saw, and does not substantively engage with my work. He doesn’t discuss any of my studies, doesn’t meaningfully engage with, much less refute, any of my objections, and in general simply makes assertions without presenting substantive arguments or evidence for his claims. This is the very definition of unsubstantiated. Go look at the tweet. Where are Rian’s arguments against my research? Where is the analysis of my data showing that his interpretation is correct and mine isn’t? You won’t find it, because it isn’t there. Rian is just making claims without supporting them. For someone who wants to lecture me about Epistemology 101, Rian appears to have a dismal understanding of what it means to present arguments, evidence, and reasons for a position. Rian hasn’t even presented much evidence that he’s read or understood any substantial portion of my work, much less that he has good reasons for his assessment of it.
After not presenting any substantive arguments or evidence for his assessment of my research, Rian then continues with his nonsensical and unsubstantiated accusations about me lying:
If you say that, sorry, then: a) either you’re being dishonest/lying - and you could write a thousand pages defending it and it would still be a lie (the part about doing it “wholeheartedly” doesn’t change the fact that you can still be lying: you know, there's such a thing as self-deception, motivated rationalization or simply maintaining two different cognitive standards at once; great news: sincerity about your conclusion doesn't imply sincerity in how you represent the reasons available for that conclusion, since you can sincerely believe that P is false and still lie about whether there is evidence for P) or b) you’re ignorant enough to not understand the distinction between: (i) there being good evidence/justification or substantive empirical support for a position and (ii) that evidence actually settling the >truth< of the position.
Merely because you’ve constructed a dichotomy doesn’t mean it’s legitimate or that the possibilities you present are the only ones on offer. What are Rian’s arguments that I’m lying? What is his evidence that I’m lying? I can’t find any. Maybe you can. And he continues to make confused remarks:
and you could write a thousand pages defending it and it would still be a lie (the part about doing it “wholeheartedly” doesn’t change the fact that you can still be lying
Rian apparently doesn’t understand the difference between actuality and possibility. Here, Rian insists I can be lying. Of course I could be lying! I’ve never disputed that it’s possible I’m lying. I’m arguing that I’m not lying, and that Rian has no good arguments, evidence, or reasons for thinking that I am. Where’s your evidence I’m lying, Rian? What are your arguments for this accusation?
Rian also misrepresents my objections to his accusations of lying. First, the conventional meaning of “lying” is to intentionally make claims you believe are false for the purpose of deceiving others. Rian’s accusations all occurred in contexts in which this is the most plausible interpretation of what he meant; at no point did Rian cancel the pragmatic implicature that would prompt this interpretation by qualifying his remarks to suggest that I may be a pathological liar or profoundly self-deceived. Given this, the bulk of my objections were directed against what a reasonable person would interpret Rian to have been claiming: that I was intentionally lying. While Rian floats the possibility that I could be “lying” in some unintentional sense, nothing about what he initially said could be reasonably interpreted in line with these possibilities.
Furthermore, is Rian accusing me of unintentional lying? If so, why not say so? And if he is: great, then my previous points are good evidence against intentional lying, which would at least pressure Rian to now make a narrower claim: not that I’m lying either intentionally or unintentionally, but instead that I’m specifically lying unintentionally.
If Rian would like to specifically accuse me of unintentionally lying, fine: what are his arguments or evidence for this? Merely floating the possibility does nothing to actually support the claim that I’m doing so.
Rian also says:
sincerity about your conclusion doesn't imply sincerity in how you represent the reasons available for that conclusion, since you can sincerely believe that P is false and still lie about whether there is evidence for P) or b) you’re ignorant enough to not understand the distinction between: (i) there being good evidence/justification or substantive empirical support for a position and (ii) that evidence actually settling the >truth< of the position
I never claimed there was no evidence for Rian’s claim; I claimed there was no good evidence for it and that it was unsubstantiated. Rian has yet to present any good reason to think I’m lying (intentionally or otherwise) about this.
Rian then presents the suggestion that I’m so ignorant I don’t understand the difference between there being good evidence/justification for something and evidence that settles the truth of that position. But this once again is very silly: I’ve never conflated these. Of course I deny there’s enough evidence to settle on the truth that moral realism is extremely intuitive across populations. That’s a given. But what I am also denying is that there’s good evidence, justification, or substantive empirical support that moral realism is widespread across populations. Rian appears to think that I think (ii) above, which presumably would be justified on his view, but conflated this with (i), and asserted (i) instead because I’m too stupid to know the difference. No. I’m just asserting something closer to (i).
It may be that Rian is so convinced that (i) is true that he can’t imagine that anyone could sincerely disagree. If so, then the mistake at the heart of this whole dispute is a simple and sad one: that with respect to certain claims, Rian may be incapable of imagining that someone could sincerely disagree with him about those claims.
Rian continues:
Those are obviously not the same thing. Evidence can make a belief rationally warranted without making it true, let alone conclusively true: that distinction is elementary epistemology.
Nobody involved in this conversation thinks otherwise. Rian is tilting at windmills. Rian makes a few remarks after this that I want to conclude on, so let’s skip ahead a bit:
Now, about the empirical research: in fact, the paragraph you quoted expressly acknowledged substantial metaethical pluralism/intrapersonal variation/context-sensitivity and the methodological dispute about whether disagreement tasks transparently measure metaethical commitments.
I’m aware of what the articles I quote say. I can quote an article to make a point without endorsing the other remarks made in the article. I don’t agree with the authors of the article that there’s substantial evidence of metaethical pluralism.
Replying that these studies do not support the claim that most people are moral realists is simply attacking a stronger thesis that I went out of my way to disclaim.
Fair enough, but this is still irrelevant. Your original claim was vague and you never clarified, so it wasn’t clear what your claim was. But either way, I don’t think moral realism is a common view among nonphilosophers, so I also reject just about any plausible instance of a weaker claim you could be endorsing. Given this, I don’t only critique the stronger claim and nothing about my position specifically centers on the notion that most people are realists. I think the actual number is negligible and that it’s just not true that it’s “extremely intuitive” to people in general, other than to at best an idiosyncratic minority.
Your work, even if it were strong (and it isn’t) challenges the construct-validity inference from certain responses to full-blown realist commitments: that is not even remotely the same claim as showing that the objectivist-seeming response pattern itself is unsubstantiated.
Rian once again claims that my work isn’t strong without any arguments, evidence, or substantiation. He just asserts this. Now, Rian’s remark is not very clear, but he seems to be saying something like that all my research does is threaten the validity of the studies that purport to show that many people choose realist response options, i.e., what my findings do is challenge whether we’re justified in inferring that these findings show that the people in question do, in fact, have realist commitments. But, Rian insists, this isn’t the same thing as showing that the mere fact that these studies show that people give “objectivist-seeming” responses is unsubstantiated.
Rian is correct. This isn’t the same thing! Which is entirely fine, because I never claimed that the latter claim was unsubstantiated in the first place. Rian is drawing a distinction I agree with, and arguing that evidence of the former isn’t evidence of the latter. The problem is that my position is only about the former and not the latter, so this is completely irrelevant. In fact, it wouldn’t even make any sense to take evidence of the former as evidence against the latter: the whole point of the former claim is that measures which purport to show that a large number of participants are realists fail to accurately measure realism. The former claim, in other words, relies on the truth of the latter claim.
For comparison, imagine I said “While it seems like that object is red, it’s actually not.” Now imagine Rian came along and said that I am lying because there’s lots of evidence that the object seems red. This would make no sense. The claim I’m making is that it isn’t red, despite seeming to be red. Not that it doesn’t even seem red. This claim presupposes agreement with the claim that the object seems red.
Rian’s other comments don’t always seem to consistently suggest he’s drawing this distinction and only arguing for the latter, but I’ll try to clear up the confusion: my position is that there’s no good evidence significant numbers of nonphilosophers have realist stances or commitments. I also claim that while many studies purport to show that this is the case, every single one of these studies relies on invalid measures, and that there is no good evidence that the participants in these studies in fact have realist stances or commitments in significant numbers.
If Rian is only arguing that there are lots of studies that purport to show that nonphilosophers are realists, then I not only don’t disagree, my entire position presupposes this is true.
If, instead, Rian is not arguing for this but is instead arguing that whatever their methodological shortcomings, these studies actually do provide so much good evidence that a significant number of nonphilosophers are realists, then I disagree, but in that case his remarks here are very weird: why make a point of saying the latter is very different from the former?
Or maybe I’ve misinterpreted Rian. I can’t tell because he’s often unclear. However, it does appear that Rian is drawing this distinction and correctly recognizes what my claim is:
Indeed, your own 2026 paper opens by saying that “most research” in experimental metaethics reports substantial interpersonal and intrapersonal variation in support for realism and antirealism and explicitly notes that earlier studies found high rates of both: let's remember that your objection is to what those responses warrant us in attributing to participants, not to whether such responses occur.
Yes. Lots of studies report high rates of moral realism. As you can see here, not only have I never denied this, I begin with this assumption in my work. The whole point of my view is that people reporting high levels of moral realism aren’t justified in interpreting these studies as evidence that these responses are good evidence that many people actually have realist intuitions because their measures aren’t valid. As I say in the abstract:
Most research in experimental metaethics suggests high levels of interpersonal and intrapersonal variation in support for moral realism and antirealism among nonphilosophers (i.e., people without significant training in philosophy). These findings challenge the assumption that nonphilosophers share a uniform commitment to moral realism. Recent evidence challenges the validity of these studies by demonstrating that many people do not interpret questions about metaethics as intended.
So what is Rian’s point? Why is Rian drawing a distinction between these claims if he’s going to “remind” me of my own position and do nothing with the distinction? This doesn’t make any sense.
Rian then returns to the Sousa et al. (2021) paper above. This return is immensely embarrassing for Rian. Rian says:
This is particularly relevant because it does not merely recycle the original disagreement task: the authors explicitly modify earlier designs to reduce alternative interpretations and then obtain high rates of objectivist responding.
Rian presents this as if what they did is modify the design of the study to reduce the kinds of alternative interpretations Moss and I raise in our 2020 paper and that I raise later in my dissertation. But this is false. This study doesn’t address the alternative interpretations we raise at all. In fact, I had contacted Sousa in 2021 in regard to the 2021 paper. From private correspondence he explicitly confirmed that he hadn’t read even our 2020 paper prior to authoring this paper. Since my dissertation hadn’t been written yet, he also couldn’t have read that. So this paper could not have addressed the specific concerns we raised explicitly in either paper, at least not explicitly and intentionally, unless they somehow anticipated precisely those concerns (which they hadn’t, and which they don’t discuss in the paper, so this possibility is moot).
The Sousa et al. (2021) paper instead only focuses on minimizing the inference that Mamilons (the Amazonian warrior tribe) and Pentars misunderstood the moral issues they were presented with. This is a narrow attempt at resolving a problem distinctive to a single study; it doesn’t solve other problems with that study and doesn’t address the general set of problems Moss and I raise at all.
Rian simply doesn’t know what he’s talking about.
In Study 2, the objectivist response was selected by 82% of participants in the same-culture condition and 74% in the exotic-culture condition. In Study 3, once the alien appraiser was explicitly described as understanding the morally relevant properties of the harmful action, objectivist responding was 74% and 70%, versus 48% when that understanding was absent.
So what? I’ve never denied these studies can and do yield high rates of realist responses. That doesn’t mean the measures used in this study are valid. They’re not. They’re just more iterations of Sarkissian et al.’s disagreement paradigm, which isn’t a valid measure. They don’t do anything to circumvent the validity issues I’ve raised. All these findings show is that when they modified the original Sarkissian et al. study’s wording it closes the gap in realist response rates across conditions. That doesn’t show that these measures are valid and that we therefore have strong evidence that many people are realists; again, and I cannot stress this enough: the modifications to this study are only specifically designed to address methodological concerns with Sarkissian et al. (2011); they are not intended to and don’t serve to address the methodological concerns myself and others have raised.
Rian continues:
Studies 4a and 4b are even more inconvenient for your blanket statement. In the injustice conditions, objective-wrongdoing responses were 88% for killing and 88% for stabbing in Study 4a, and 92% for killing and 81% for stabbing in Study 4b. Study 4b deliberately used the much more explicit formulation that the action was “inherently wrong,” i.e. wrong independent of any prevailing cultural norms. The authors themselves describe Studies 2–3 as showing “high objectivist responding” and Studies 4a–4b as showing a “quite high rate” of objective-wrongdoing responses.
No, none of this is remotely inconvenient for my position. That the realist response rates are higher here isn’t good evidence of the validity of the measures. It would hardly matter if these studies found 100% realist response rates. As long as there are good reasons to question the validity of these measures that the authors haven’t addressed (and there are), and as long as there’s no adequate evidence of their general validity, then there’s no good reason to trust these measures any more than other measures that I’ve critiqued. 100% of people can interpret stimuli in unintended ways, just as 85% or 50% of people can.
This remark in particular is rather silly:
Study 4b deliberately used the much more explicit formulation that the action was “inherently wrong,” i.e. wrong independent of any prevailing cultural norms.
It’s odd for Rian to drop the full quote and present the latter description outside the quotes. What the item states is that the action was:
inherently wrong, that is, it is wrong independent of any prevailing cultural norms."
Now, this is a problem for a few reasons. First, whether something is wrong independent of cultural norms only technically rules out cultural relativism; it does not rule out other forms of antirealism. So this already isn’t a good way to measure whether a person is a realist or an antirealist since it measures whether a person is a cultural relativist vs. whether they’re not a cultural relativist. Rejecting cultural relativism does not mean you’re a realist. I reject cultural relativism. I’m not a realist.
Second, do we have any reason to think that this explicit formulation of “inherently wrong” is an improvement and resulted in a more valid measure? I don’t think so, and Rian doesn’t provide any. Generally speaking, explicit formulations can be an impediment because they can prompt metalinguistic theorizing and actually interfere with a valid measure, so there’s even reason to think that explicit formulations are worse in some cases. This may be how Rian interprets that phrase. But what matters is how participants interpreted that phrase. That participants may interpret stimuli including phrases like this in unintended ways is the main point of my critique, and my dissertation research shows that intended interpretation rates are exceptionally low for every tested measure, including the terminology used in Fisher et al.’s (2017) explicit formulation that used the term “objective.”
So what we have here is a novel measure that might be interpreted as intended, but for which there is little supporting evidence. In fact, I wasn’t immediately aware of what study Rian was referring to when he mentioned Sousa. This is because I assumed he was referring to a newer study. Why refer to an old study that doesn’t actually make any significant methodological headway over prior studies? But Rian was referring to this study, which was published five years ago, before I had completed my dissertation, and that doesn’t do or even attempt to do what Rian says it does.
Sousa et al. gathered open response data that provides some insight into how participants may have interpreted the stimuli in this study. However, they do not offer any systematic examination of this data. Instead, they only report a handful of select examples from Study 4a (and none from 4b, which are the stimuli Rian mentions above):
we believe most participants choosing “Person A is correct” in fact expressed an objectivist position. At least this interpretation was often conveyed by their justifications:
“I believe that there is such a thing as moral truths. (…)” (Killing)
“Person A is right, just because Baako doesn’t have the moral intelligence to know that the stabbing is wrong does not make it okay.” (Stabbing)
Neither of these responses provides clear or even good evidence that participants interpreted the prompt as intended. First, they cut off the first of these quotes. This first quote is from response #181 in their raw dataset. Here is the full quote:
I believe that there is such a thing as moral truths. I think that value systems vary from better to worse depending primarily on how well they promote or are an obstacle to the flourishing of sentient beings. Killing children for being unattractive sees [sic] to bad for human flourishing.
Notice the latter remark: that value systems can vary from better to worse. Does this sound like something a moral realist would say? No, it does not. This might instead be a judgment about how closely first-order normative systems accord with the respondent’s own normative moral standards, in which case this response would indicate an unintended interpretation. Alternatively, if this remark is intended as metaethical, it could reflect at least in part a somewhat more relativistic view about morality.
The second remark likewise does not indicate that the participant interpreted the stimuli in line with a realist interpretation. It sounds like they are saying that the person from the other culture has compromised cognitive faculties associated with moral judgment; one might imagine a person with psychopathy or who was brainwashed into thinking violence or murder are okay. This person appears to be saying that just because a person with such deficiencies thinks such actions are okay, this does not make them okay.
This does not indicate that the participant is taking a realist stance. On the contrary, a more conventional realist interpretation actually conflicts with these remarks. The best indicator of realism would be where one fully appreciates that the other person is in possession of a capacity of reasoning about morality, but has reached a contrary conclusion to you, and yet you still maintain that they’re mistaken. This participant appears to be offering an explanation for the psychology of the other person that exempts them from being a genuine peer with a contrary moral perspective. In other words, rather than considering their disagreement with this person to be a disagreement with a fully cognitively capable moral agent who has arrived at a different conclusion about morality, they appear to be inferring that the person lacks moral competence, much as we would not consider the actions of an insane person or an animal to be morally permissible merely because the insane person or animal doesn’t have the moral faculties to appreciate the moral status of their actions.
In other words, the two examples Sousa et al. present as evidence that participants interpreted their stimuli as intended not only fail to do so, they might even serve as evidence to the contrary! And keep in mind they could have chosen any examples from the data to support their claim; if this is the best they can do, this should if anything lead us to be suspicious that any appreciable number of participants genuinely appear to have interpreted the stimuli as intended.
Unfortunately, they offer no systematic evaluation of their open response data that would support the conclusion that most participants interpreted the prompt as intended (i.e., in line with realism). Incidentally, I had commented on this fact in my dissertation:
Even when researchers do appeal to open response data, they rarely do so in a systematic and thorough way. For instance, Sousa et al. (2021) appeal to open response data to support their explanation of their findings. However, they only highlight one or two examples to support a point, rather than appealing to systematic analysis of the data they collected. (p. 101, footnote 61)
As it happens, I had the raw data from their open response questions at the time and have already examined it. Their data is publicly available here and you can go see it for yourself. If you look at the open responses for Study 4a, they don’t show a massive shift towards intended interpretation rates relative to other studies.
Hurting people is wrong, even if a culture determines it is right.
It is wrong and it breaks a moral rule - murder is wrong - that all societies should follow.
Just because one society does it doesn’t mean all societies live by that rule.
I agree with Baako
Person A is letting his or her flawed view of society interfere with making the right decision. Kill the sick child is a simple way to eliminate it’s suffering and protect the whole tribe from being burdened, a sensible move.
Stabbing should never be permissable.
Stabbing a random person is totoally reckless and unnecessary.
stabbing someone else should simply not be acceptable. no matter who you are.
I feel that it is considered murder to kill the healthy child. So I would agree with person A.
It is morally wrong to kill anyone
It’s wrong to kill anyone unless they killed someone else.
The killing of someone innocent that is not for the greater good is wrong regardless of beliefs
Maybe it is a good idea. Who is Person A to say if this is bad. Maybe they make sure to not stab in a deadly way.
A is employing a discourse, i.e., moral discourse, which is a subjective expression of will to power. Moral claims have no universal intersubjective validity even when applied intraculturally.
It’s not wrong if that’s what they believe
Person A should respect the cultural diversity and norms that are present in Mamilon society. It is not his place to judge them as an outsider.
Because taking a life for something like that is wrong
Person A is correct within the constraints of person A’s moral code and agreement of societal norms.
Hurting an innocent person in any way is wrong morally and generally.
Let’s go through these. The first is the closest to a good sign of intended interpretation and if I were coding the responses here I’d label this as an intended interpretation. But note that this is already being extremely generous. Such a remark potentially illustrates only a superficial capacity to recapitulate information provided in the instructions or stimuli. After all, participants are told that the person thinks “What Baako did is still wrong, even if the Mamilons do not think it is wrong.” It is not difficult for a person who has little or no understanding of relativism or antirealism to respond with a first-order moral judgment, and then dismiss the idea that a culture could fix the rightness of an action, without understanding the sense in which relativists think cultures serve to determine the truth of moral claims.
Skepticism about the superficiality of such responses is warranted given both the overarching data I have on participant interpretation, and David Moss’s experience with interviewing students about their metaethical views (Bush & Moss, 2016). Even so, proponents of validity are off to a good start with this remark. Subsequent remarks quickly reveal that this first response was an anomaly. The second response expresses a normative standard and universalism, not anything related to realism (universalism is not the same thing as realism). The third makes a descriptive claim. The fourth simply expresses agreement, without any explicit metaethical content. The fifth judges the person and offers a reason for why they might make a given choice, but doesn’t clearly frame this in metaethical terms. The sixth expresses an absolutist view rather than a view about realism or antirealism, and so on.
As you can see, there’s no clear pattern of responses indicating a propensity to draw on metaethical considerations. The 12th item hints at a metaethical view when it says “regardless of beliefs,” the 14th conveys fairly sophisticated metaethical views, and the 15th appears to endorse agent relativism. A generous reading of these responses would indicate maybe about 20% say things that support an intended interpretation. This is in line with the data reported in my dissertation. Again, though, note that a person could have a metaethical interpretation but not indicate it in their response, or appear to have a metaethical interpretation but have an inadequate or superficial understanding of the stimuli, so findings like these are not an especially compelling bit of evidence on their own. Most importantly, this particular line of questioning simply asked the participant to explain their response and did not specifically probe their interpretation, so it’s not well suited to addressing interpretation rates.
However, Rian emphasized the stimuli used in Study 4b. What do the authors say about it?
More importantly, we replicated this result in Study 4b, in which the response probe differed from that in 4a, such that it articulated a more unequivocally objectivist position (“Baako’s action is inherently wrong, that is, it is wrong independent of any prevailing cultural norms.”). Neither a relativist, a nihilist nor a subjectivist could choose the option “Person A is correct” here.
This isn’t even true. Notice that they define “inherently wrong” as wrong in a way that is “independent of any prevailing cultural norms.” “Inherently wrong” is a bit of unclear jargon. An untrained participant with no background in philosophy could interpret it in a variety of ways. But the only way for it to be interpreted in a way that relativists, nihilists, and subjectivists couldn’t choose option A is if they understood it to mean something like “stance-independently true,” and, technically, since relativism and subjectivism are consistent with stance-independence, it would have to be stance-independent and we’d have to ignore realist forms of relativism or presume they also inferred universality. But we can set all of that aside because the stimuli instruct readers in how to interpret “inherently wrong”: to mean wrong in a way that isn’t determined by culture. Can relativists, nihilists, and subjectivists judge that an action is wrong independent of prevailing cultural norms?
Yes. Yes they can. Individual relativists wouldn’t think cultural norms determine right or wrong, nor would quasi-realists, expressivists, constructivists, fictionalist error theorists, quietists like myself, and so on. In fact most antirealist positions are consistent with this, besides traditional noncognitivism, cultural relativism, and non-fictionalist error theory. So some antirealists couldn’t consistently select this option, but many could.
Note, then, that the authors themselves are mistaken about the responses antirealists could consistently offer in response to this prompt. However, this is largely moot, because this is all only relevant on the assumption participants have these views and have interpreted stimuli as intended. In other words, their claims about how a person could or couldn’t respond require that participants have these views, require that they interpret the prompt as intended, require that they somehow circumvent the misleading information that the authors provide that qualifies “inherently wrong” to mean only culture-independent when it should have been stance-independent in a more general sense, and require that they answer in a way consistent with their positions. Those are a lot of requirements and notably they presuppose precisely those assumptions that I am denying: that participants do hold these views and that they do interpret the stimuli as intended. The authors’ claims are thus not an especially good reason to think the stimuli serve as a valid measure. They’re not even intended to; this isn’t a claim about validity in a direct sense; it’s a claim about the accuracy of their operationalization in principle. The responses to this prompt aren’t any more impressive than for 4a.
Overall for Study 4b you get the same pattern of people repeating the instructions from the prompts/stimuli, expressing their first-order moral views, and so on. Participants do not appear to consistently respond in a way that suggests they interpreted the question in metaethical terms. Does this mean they didn’t interpret it in metaethical terms? No, but again, given that we have so many other reasons to doubt the validity of metaethics measures, and especially the disagreement paradigm, this coupled with the same pattern in their open response data is most consistent with my stance, which is that we expect this measure not to be valid, either.
Many participants simply reiterate that it’s wrong to kill or insist it’s wrong to kill “for any reason.” Such responses remain in the realm of first-order judgment and at best suggest participants are judging the action in terms of whether there are exceptions to norms against killing. Many others say things that are ambiguous or unclear and don’t tell you one way or another whether they interpreted the stimuli as intended. You do get a handful of people that do respond in ways that either strongly hint at or quite explicitly suggest an intended interpretation. For instance, you get responses like this:
Morality is relative and depends on cultural norms
No matter what people think, it is horrible to kill a healthy being, whether it be human, animal, or beast.
These are good responses, and the kind I’d code as indications of intended interpretation. These sorts of responses are quite rare, though, and probably comprise less than 20-30% of responses. Note that not giving a response like this doesn’t mean that participants didn’t interpret stimuli as intended. Intended response rates might be higher. Conversely, offering a response like these isn’t strong evidence that they do interpret stimuli as intended, since these remarks are often shallow and underdeveloped, or invoke technical terminology that other data suggests nonphilosophers frequently understand in ways that are inconsistent with the concept being studied by researchers. This kind of open response data just isn’t very diagnostic one way or another. This is due in part to what Sousa et al. asked. They asked participants to explain why they answered the question in the way they had. This is a good question to ask, and can reveal important data about how participants reason, but it’s not specifically designed or optimized for assessing intended interpretation. Given this, we shouldn’t put much stock in this data as evidence for or against the validity of the measures. But from a cursory glance, it looks to me to provide some small evidence against validity.
Rian also continues to push an irrelevant distinction:
And notice the distinction I made from the beginning: Sousa et al. themselves explicitly caution that these results do not settle whether folk objectivism amounts to correspondence-style moral realism.
Of course they don’t settle it. Rian continues to be confused about what my position is. None of us working in the field consider individual studies like these to be decisive or to settle anything. That’s never been at issue. I’m not confused about this. Rian seems to think, for no apparent reason, that I’m confused about the difference between evidence decisively establishing something and it justifying it. I’m not confused about this distinction and never drew on or even alluded to it. I’m not merely saying that studies like these don’t confirm or settle the question of whether realism is common among nonphilosophers; I’m saying these studies don’t provide good enough evidence at all to justify thinking that realism is common among nonphilosophers. This is a different interpretation of the results of these studies than those presented by the authors themselves; it’s as if Rian somehow thinks it’s not possible or defensible for me to disagree with the authors of a study, or if the authors say something about their findings, that it must be true.
I’m not sure why Rian thinks I’m confused about this. I meant exactly what I said. Maybe Rian’s incredulity that someone could take this strong of a stance has led him to think it’d only be charitable to assume I couldn’t really be making this strong of a claim. Rian is wrong. I am making that strong of a claim. I am saying there is not a single study in existence, nor is there a credible body of literature as a whole that justifies belief that realism is common among nonphilosophers. I’m not just saying we don’t have decisive evidence of widespread realism. I’m saying we don’t have any good evidence at all, including this study.
Rian continues:
Objectivist intuitions “might be realist,” they say, but might instead admit another interpretation. Exactly. That is why I said objectivist or realist-seeming, rather than claiming that these experiments establish that most people are card-carrying metaethical realists!
Rian continues to be confused. The authors take their findings to be based on valid measures that provide significant but non-decisive justification for moving in the direction of the conclusion that many nonphilosophers are moral realists about at least some issues. That their findings move the dial in this direction would only be true if their measures were valid. They think they are. A single set of studies presenting good evidence of a claim is still interpreted among scientists as tentative because we care about what the overarching body of literature as a whole ultimately says about the matter. Yet Rian goes back to using the ambiguous phrase “realist-seeming,” and contrasts this with the notion that the evidence establishes their hypothesis. Again, Rian is mixing things up. There are four claims in play here:
Whether lots of participants choose the realist responses in these studies.
Whether the measures used in these studies are valid, and that we therefore have at least some significant evidence that many nonphilosophers are moral realists.
Whether the measures used in these studies are valid, and the evidence is so strong so as to justify belief that many nonphilosophers are moral realists.
Whether the results of the study are overwhelming and decisive evidence that many nonphilosophers are moral realists.
Rian seems to think I’m mixing these up, and mistakenly thinks that I am only making the claims I do because I think the alternative is (4). I don’t think this. I take a holistic approach to bodies of literature. We don’t typically judge whether a hypothesis is true or false on the basis of a single study. We look at the literature as a whole. Any one study could appear in isolation to be very strong evidence for a claim. But this evidence must be weighed against all the other evidence and theoretical considerations to the contrary. This approach to evaluating evidence is so standard that it operates in the background among scientists.
That Rian would so blithely presume I think otherwise, and get tangled up in such elementary confusions and assumptions about how I and others would think about data is little more than a testament to Rian’s own ignorance and a strong indication that Rian is an amateur with little understanding of how experienced researchers evaluate evidence. Mind you: Rian presents no good reasons at all to suppose that I think this way. He just assumes it, apparently based on his own misguided assumptions. I’ve made numerous public remarks about holistic evaluation of the evidence and routinely speak about the importance of converging lines of evidence. Nobody familiar with my past statements or how I think would make these errors. Rian is unfamiliar with the literature, unfamiliar with where the Sousa et al. paper fits into the literature, unfamiliar with my past statements and ways of thinking, and unfamiliar with conventions and norms about how scientists reason and think about evidence. Rian appears to be so thoroughly ignorant of just about everything it is remarkable that he appears so confident. But maybe it shouldn’t be so remarkable. Some people are so profoundly misinformed that they aren’t aware of how badly misinformed they are. I think this is what’s going on with Rian.
Rian concludes:
Sorry, but based on everything so far, it seems to me that the inclusive disjunction I laid out there about you stands still.
In other words, Rian thinks what he’s just said supports the conclusion that I am lying and/or dumb. Let’s return now to a remark Rian made earlier:
Applying the principle of charity to your intelligence, I initially went with the first option: that you know the distinction and are simply being dishonest about what the literature actually supports. But the second option remains entirely possible as well. And, in case you missed it: this is an inclusive disjunction, since both a) and b) could be true. So here's my retraction: Lance, you're lying or being ignorant/dumb (maybe both).
No, Rian. I disagree with you about what the literature supports. You have presented no good reasons or evidence to conclude that I am lying or that I’m being ignorant or dumb.
Rian’s suggestion that I am dumb and/or lying is ironic given that Rian has, time and again, displayed an astonishing inability to understand me and to understand the experimental metaethics literature. He shows no indication of being in a credible position to judge my intelligence or my honesty.
Such profound confusion would be unfortunate but hardly a vice were it not coupled with a cartoonish degree of childish smugness, condescension, and unwarranted hostility. This is not the behavior of a serious or thoughtful person. It is the behavior of an angry, petulant, and immature person who has a lot of growing up and a lot of reading to do.
References
Baguley, T. (2012, June 21). The stimuli-as-fixed-effect fallacy. R-bloggers. https://www.r-bloggers.com/2012/06/the-stimuli-as-a-fixed-effect-fallacy/
Beebe, J. (2015). The empirical study of folk metaethics. Etyka, 50, 11-28.
Beebe, J., Qiaoan, R., Wysocki, T., & Endara, M. A. (2015). Moral objectivism in cross-cultural perspective. Journal of Cognition and Culture, 15(3-4), 386-401.
Bush, L. S. (2023). Schrödinger's Categories: The Indeterminacy of Folk Metaethics. Cornell University.
Bush, L. S. (2026). Why training paradigms won’t rescue experimental metaethics. Review of Philosophy and Psychology, 1-26.
Bush, L. S., & Moss, D. (2016, September). Qualitative studies in metaethics: Variability, inconsistency, & indeterminacy. Invited talk at the 2016 Buffalo Annual Experimental Philosophy Conference, Buffalo, NY.
Bush, L. S., & Moss, D. (2020). Misunderstanding metaethics: Difficulties measuring folk objectivism and relativism. Diametros 17(64): 6-21.
Davis, T. (2021). Beyond objectivism: New methods for studying metaethical intuitions. Philosophical Psychology, 34(1), 125-153.
Fisher, M., Knobe, J., Strickland, B., & Keil, F. C. (2017). The influence of social interaction on intuitions of objectivity and subjectivity. Cognitive Science, 41(4), 1119-1134.
Flake, J. K., Pek, J., & Hehman, E. (2017). Construct validation in social and personality research: Current practice and recommendations. Social Psychological and Personality Science, 8(4), 370-378.
Goodwin, G. P., & Darley, J. M. (2008). The psychology of meta-ethics: Exploring objectivism. Cognition, 106(3), 1339-1366.
Judd, C. M., Westfall, J., & Kenny, D. A. (2012). Treating stimuli as a random factor in social psychology: A new and comprehensive solution to a pervasive but largely ignored problem. Journal of Personality and Social Psychology, 103(1), 54–69. https://doi.org/10.1037/a0028347
Moss, D., Montealegre, A., Bush, L. S., Caviola, L., & Pizarro, D. (2025). Signaling (in)tolerance: The reputational implications of metaethical relativism and objectivism. https://doi.org/10.31234/osf.io/62z8c
Nichols, S. (2002). Norms with feeling: Towards a psychological account of moral judgment. Cognition, 84(2), 221-236.
Nichols, S., & Folds-Bennett, T. (2003). Are children moral objectivists? Children's judgments about moral and response-dependent properties. Cognition, 90(2), B23-B32.
Pölzler, T. (2017). Revisiting folk moral realism. Review of Philosophy and Psychology, 8(2), 455-476.
Pölzler, T., & Wright, J. C. (2020). Anti-realist pluralism: A new approach to folk metaethics. Review of Philosophy and Psychology, 11(1), 53-82.
Reinecke, M. G., & Solomon, L. H. (2023). Children deny that God could change morality. Cognitive Development, 68, 101393.
Sarkissian, H., Park, J., Tien, D., Wright, J. C., & Knobe, J. (2011). Folk moral relativism. Mind & Language, 26(4), 482-505.
Sousa, P., Allard, A., Piazza, J., & Goodwin, G. P. (2021). Folk moral objectivism: The case of harmful actions. Frontiers in Psychology, 12, 638515.
Wright, J. C., Grandjean, P. T., & McWhite, C. B. (2013). The meta-ethical grounding of our moral beliefs: Evidence for meta-ethical pluralism. Philosophical Psychology, 26(3), 336-361.
Zijlstra, L. (2023). Are people implicitly moral objectivists? Are People Implicitly Moral Objectivists?. Review of Philosophy and Psychology, 14(1), 229-247.






