No abstract is provided for this article.
Theories of measurement are of central importance to test score interpretations in terms of measurement. Psychometric models are of central importance to the
The question of whether psychopathology constructs are discrete kinds or continuous dimensions represents an important issue in clinical psychology and psychiatry. The present paper reviews psychometric modelling approaches that can be used to investigate this question through the application of statistical models. The relation between constructs and indicator variables in models with categorical and continuous latent variables is discussed, as are techniques specifically designed to address the distinction between latent categories as opposed to continua (taxometrics). In addition, we examine latent variable models that allow latent structures to have both continuous and categorical characteristics, such as factor mixture models and grade-of-membership models. Finally, we discuss recent alternative approaches based on network analysis and dynamical systems theory, which entail that the structure of constructs may be continuous for some individuals but categorical for others. Our evaluation of the psychometric literature shows that the kinds–continua distinction is considerably more subtle than is often presupposed in research; in particular, the hypotheses of kinds and continua are not mutually exclusive or exhaustive. We discuss opportunities to go beyond current research on the issue by using dynamical systems models, intra-individual time series and experimental manipulations.
One of the key concepts in the research on political attitudes is attitude strength. Strong attitudes are durable and impactful, while weak attitudes are fluctuating and inconsequential. Recently, the Causal Attitude Network (CAN) model was proposed as a comprehensive measurement model of attitudes. In this model, attitudes are conceptualized as networks of causally connected evaluative reactions (i.e., beliefs, feelings, and behavior toward an attitude object). Here, we test the central postulate of the CAN model that strong attitudes correspond to highly connected attitude networks. We use data from the American National Election Studies 1980-2012 on attitudes toward presidential candidates (total n = 18,795). The results show that attitude strength and connectivity of attitude networks are strongly related. Additional analyses show that connections between non-behavioral evaluative reactions (i.e., beliefs and feelings toward presidential candidates) are highly predictive of the connections between behavior (i.e., voting decisions) and non-behavioral evaluative reactions. This result indicates that connectivity of political attitude networks accounts for differences between strong and weak attitudes in attitude-behavior consistency with respect to voting decisions. We conclude that network theory provides a promising framework to advance the understanding of attitude strength.
Test equating under the NEAT design is, at best, a necessary evil. At bottom, the procedure aims to reach a conclusion on what a tested person would have done, if he or she were administered a set ...
The Ising model is one of the most popular models in network psychometrics. However, statistical analysis of the Ising model is difficult due to the presence of its intractable normalizing constant in the probability function. As a result, maximum likelihood estimation using the exact likelihood is only possible for small graphs, and approximation methods are needed for larger graphs. Two popular approximations of the exact likelihood are the joint pseudolikelihood (JPL) and the disjoint pseudolikelihood (DPL). These approximations yield consistent estimators, but we do not know how well they perform for finite data. In this paper, we investigate the relative performance of parameter estimation methods based on the two approximations and compare them to maximum likelihood estimation using the exact likelihood. We perform an extensive simulation study comparing the estimators in terms of bias and variance. We show that maximum pseudolikelihood estimation based on the JPL is a stable estimation method that is able to accurately approximate the maximum likelihood estimates, but that maximum pseudolikelihood estimation based on the DPL only works well for large sample sizes.
Items in a test are often used as a basis for making decisions and such tests are therefore required to have good psychometric properties, like unidimensionality. In many cases the sum score is used in combination with a threshold to decide between pass or fail, for instance. Here we consider whether such a decision function is appropriate, without a latent variable model, and which properties of a decision function are desirable. We consider reliability (stability) of the decision function, i.e., does the decision change upon perturbations, or changes in a fraction of the outcomes of the items (measurement error). We are concerned with questions of whether the sum score is the best way to aggregate the items, and if so why. We use ideas from test theory, social choice theory, graphical models, computer science and probability theory to answer these questions. We conclude that a weighted sum score has desirable properties that (i) fit with test theory and is observable (similar to a condition like conditional association), (ii) has the property that a decision is stable (reliable), and (iii) satisfies Rousseau's criterion that the input should match the decision. We use Fourier analysis of Boolean functions to investigate whether a decision function is stable and to figure out which (set of) items has proportionally too large an influence on the decision. To apply these techniques we invoke ideas from graphical models and use a pseudo-likelihood factorisation of the probability distribution.
The concept of measurement plays an important role in validity theory. But what does the term ‘measurement’ mean? This chapter discusses theories that provide
About five decades ago, the visionary Dutch psychologist A. D. De Groot started building an extraordinary academic group at the University of Amsterdam. It consisted of psychometricians, statisticians, philosophers of science, and psychologists with a general methodological orientation. The idea was to approach methodological problems in psychology from the various angles these different specialists brought to the subject matter. By triangulating their viewpoints, methodological problems were to be clarified, pinpointed, and solved. This idea is in several respects the basis for this book. At an intellectual level, the research reported here is carried out exactly along the lines De Groot envisaged, because it applies insights from psychology, philosophy of science, and psychometrics to the problem of psychological measurement. At a more practical level, I think that, if De Groot had not founded this group, the book now before you would not have existed. For the people in the psychological methods department both sparked my interests in psychometrics and philosophy of science, and provided me with the opportunity to start out on the research that is the basis for this book. Hence, I thank De Groot for his vision, and the people in the psychological methods group for creating such a great intellectual atmosphere.
According to the correspondence theory of truth, a proposition is true if and only if the world is as the proposition says it is. This theory has been both promoted and rejected by philosophers and scientists down through time. In this paper, we adopt the correspondence theory as a plausible theory of truth and relate it to science. First, we briefly outline the major extant theories of truth. We then present the correspondence theory in a form that enables us to show that the theory uniquely fulfills a crucial function in psychological research, because the interpretation of truth claims as suppositions that concern states of affairs in the world clearly explicates what it means for a theory to be true, and what it means for a theory to be false. For this reason, correspondence truth has the advantage of allowing researchers to properly understand the assumptions of scientific research as claims about the factual state of the world, and to scrutinize these assumptions. It is concluded that correspondence truth plays an important part in our understanding of science, including psychology.
Stevens’ theory of admissible statistics [Stevens, S. S. (1946). On the theory of scales of measurement. Science, 103, 677680] states that measurement levels should guide the choice of statistical test, such that the truth value of statements based on a statistical analysis remains invariant under admissible transformations of the data. Lord [Lord, F. M. (1953). On the statistical treatment of football numbers. American Psychologist, 8, 750–751] challenged this theory. In a thought experiment, a parametric test is performed on football numbers (identifying players: a nominal representation) to decide whether a sample from the machine issuing these numbers should be considered non-random. This is an apparently illegal test, since its outcomes are not invariant under admissible transformations for the nominal measurement level. Nevertheless, it results in a sensible conclusion: the number-issuing machine was tampered with. In the ensuing measurement-statistics debate Lord’s contribution has been influential, but has also led to much confusion. The present aim is to show that the thought experiment contains a serious flaw. First it is shown that the implicit assumption that the numbers are nominal is false. This disqualifies Lord’s argument as a valid counterexample to Stevens’ dictum. Second, it is argued that the football numbers do not represent just the nominal property of non-identity of the players; they also represent the amount of bias in the machine. It is a question about this property–not a property that relates to the identity of the football players–that the statistical test is concerned with. Therefore, only this property is relevant to Lord’s argument. We argue that the level of bias in the machine, indicated by the population mean, conforms to a bisymmetric structure, which means that it lies on an interval scale. In this light, Lord’s thought experiment–interpreted by many as a problematic counterexample to Stevens’ theory of admissible statistics–conforms perfectly to Stevens’ dictum.
In World Psychiatry , het periodieke tijdschrift van de World Psychiatric Association, verscheen begin dit jaar een opmerkelijk artikel van onderzoeker Den
The usage of psychological networks that conceptualize behavior as a complex interplay of psychological and other components has gained increasing popularity in various research fields. While prior publications have tackled the topics of estimating and interpreting such networks, little work has been conducted to check how accurate (i.e., prone to sampling variation) networks are estimated, and how stable (i.e., interpretation remains similar with less observations) inferences from the network structure (such as centrality indices) are. In this tutorial paper, we aim to introduce the reader to this field and tackle the problem of accuracy under sampling variation. We first introduce the current state-of-the-art of network estimation. Second, we provide a rationale why researchers should investigate the accuracy of psychological networks. Third, we describe how bootstrap routines can be used to (A) assess the accuracy of estimated network connections, (B) investigate the stability of centrality indices, and (C) test whether network connections and centrality estimates for different variables differ from each other. We introduce two novel statistical methods: for (B) the correlation stability coefficient, and for (C) the bootstrapped difference test for edge-weights and centrality indices. We conducted and present simulation studies to assess the performance of both methods. Finally, we developed the free R-package bootnet that allows for estimating psychological networks in a generalized framework in addition to the proposed bootstrap methods. We showcase bootnet in a tutorial, accompanied by R syntax, in which we analyze a dataset of 359 women with posttraumatic stress disorder available online.