Psychometric testing

Topic
At a glance

Psychometric tests (or assessments, or testing) are tests that are carefully designed, constructed, and validated by expert psychologists and specialists to reliably measure very specific aspects of our psychological make-up. These aspects include our personality, our problem solving ability (also known as reasoning ability, or cognitive ability), critical thinking, emotional intelligence, attitudes and behaviours, and other core abilities and behavioural profiles.

Psychometric testing
In-depth

Psychometric testing, or psychological assessments, are tools that are designed and implemented to provide standardised, objective, and evidence-based (i.e. data-based) measurements of critical human psychological capabilities and elements. Key and common examples include our raw problem solving and reasoning power (or cognitive ability), personality, emotional intelligence, safety behaviours and attitudes, other behavioural and attitudinal constructs and outcomes, and other areas. Psychometric testing has emerged now as one of most effective, cost-conscious, and generalisable ways to measure key aspects of intelligence/problem-solving, likely future behaviour, and performance on given tasks, for use with job candidates in recruitment and selection, in developmental processes, and through surveying teams and organisations to gauge engagement and make more well-informed decisions in general. The core concept of psychometrics is all in the name:

  • psych - the mind
  • metric - to measure

A key defining feature of psychometric testing is the expertise in domain and skill that is exercised in the development and validation of these tools, which stands as the key determinant for the effectiveness and accuracy of a tool. Psychometrics experts (psychometricians) plan, design, and test these tools through several development stages, ensuring that the test questions are consistent with each other internally (internal consistency/reliability), that they only measure one thing, and that one thing is measured accurately (validity/construct validity), that the tool is statistically comparable to a similar tool (convergent validity), and that it predicts the real world outcomes that it is claimed to predict (criterion validity). These stipulations are challenging to achieve, and are subject to continuous improvement, iterations, and adjustments over time.

Like the broader psychological research body and literature, psychometrics as a concept and practical tool emerged from its scientific dark ages at the turn of the last century, with key figures such as Francis Galton, Wilhelm Wundt, L. L. Thurstone, Alfred Binet, Theodore Simon, and later Raymond Cattell and others. Reliance on concepts like phrenology, and pervasive erroneous beliefs based on biases and race, made way for a more nuanced and scientifically informed approach to measuring key constructs, initially for human intelligence. This manifested in foundational tools such as Army Alpha and Army Beta tests in WWI era America, developed by Robert Yerkes, and the Stanford-Binet IQ Test.

Over time more sophisticated statistical analysis techniques were developed, with approaches like factor analysis effectively leveraging test data to refine measurement and prediction. Advances in technology and computing enabled the practice of psychometric testing even further, allowing for a more standardised delivery method and automated scoring and facilitating the move away from cumbersome pencil and paper style delivery, which had been the dominant (and only available) method to administer testing to that point.

The proliferation of personal computers and powerful personal electronic devices has enabled the widespread adoption of unproctored (unsupervised) psychometric testing on the candidates' and test-takers' personal devices. It would be fair to say that this has been a double edged sword, where we potentially sacrifice the consistency for how the test is administered, and increase the potential for technical errors and dishonest test-taking practices (cheating, and cheating adjacent behaviours), we have also increased test-taker accessibility, scalability of administration, and enabled widespread adoption of this powerful and predictive tool, normalising its use and emerging as a competitive advantage when implemented effectively.

Psychometric testing now features prominently across sectors, worldwide, in business, government, education, medicine, and elite sport, these can include:

  • reasoning and problem solving (cognitive ability)
  • personality (generally a trait-based Five Factors oriented test for selection purposes, or a type-based assessment for team profiling and coaching work)
  • emotional intelligence
  • safety attitudes, and safety behaviour 
  • leadership behaviours and 360 assessments
  • surveys and team-wide assessment tools when designed with specific outcomes in mind (engagement, change readiness, psychosocial hazards, etc) 

Spotting the real deal

Psychometric assessments, on the surface, seem relatively simple to understand. So much so that many people believe that they are simple enough to put together themselves as well. This is an incorrect assumption. This has led to a vast quantity of examples of 'online tests' that look interesting, fun, intuitively 'real' and legitimate as a measure, but owing to the developers of these online 'experiences' not knowing what they don't know, are useless as a predictive tool at best, and dangerously being used to make real decisions at worst.

Validity & Reliability

The strength and capability of a psychometric test is partly determined by its measures of validity and reliability. A tests validity (via construct, convergent, and/or criterion validity) indicates whether a test is measuring what we say it is measuring, and predicting what we are saying it is predicting. A tests reliability refers to the consistency of the tool, whether its items (questions) are roughly consistent with each other (when they should be) and consistent in the responses that it elicits from test-takers from one completion to the next (test-retest reliability). One tip to keep an eye out for: while longer tests with seemingly a bunch of questions that repeat themselves can be frustrating/boring/time-consuming, and shorter and 'punchier' tests can be less arduous and seem more intuitively 'enough', including more questions is often the best and only real way to improve and strengthen both reliability and validity at the same time. If a test is quite short, this can be a trap and worth looking into further. 

Normative Comparison Groups (norms)

The 'raw' output, or raw score, of a test tends not to mean very much at all by itself. To make those scores meaningful in any way, they are taken and compared to the average scores of a much larger (ideally) group of test-takers who have completed this exact assessment already. Those group statistics can then be used to convert the single test-taker's score/s into classification bands (average, below average, above average etc) in simpler cases, or percentile ranks (percentage of that comparison group that the test-taker's scores exceed) where more granularity is required. The norm group dictates an enormous amount about how we interpret a test score, and gives it a very large proportion of it's meaning. If you have found or taken a test online and you cannot seem to find the information regarding the comparison group/s that are being used to generate your score, we would recommend prioritising finding this information or enquiring with the test provider until you have provided with this information.

Research findings

Coined the term "Big Five" to describe core personality dimensions and later created the open-source International Personality Item Pool (IPIP, 1999) (Goldberg, L. R., 1990)

Pivotal work in the modern psychometric world with the personality framework and term 'Big Five' created to describe our modern scientific understanding of the core personality dimensions that each human being features in, with this work also contributing to the International Personality Item Pool (IPIP, 1999), a critical resource for researchers and students (Goldberg, L. R., 1990). 

https://ipip.ori.org/

https://psycnet.apa.org/doi/10.1037/0022-3514.59.6.1216 

https://psycnet.apa.org/record/1991-09869-001 

Empirically validated and coined the Five Factor Model (FFM) conceptualisation of the Big Five, contemporaneous with Goldberg's Big Five (Costa, P. T., & McCrae, R. R. 1992)

Empirically validated the Five-Factor Model (FFM) across demographics and developed the industry-standard NEO-PI-R assessment instrument, which contributed to the popularisation of this form of testing with its portability and standardised administration (Costa, P. T., & McCrae, R. R. 1992)

https://psycnet.apa.org/doi/10.4135/9781849200479.n9

https://psycnet.apa.org/record/2008-14475-009 

Pivotal research linking personality to workplace performance (Barrick, M. R., & Mount, M. K. 1991)

Demonstrated via landmark meta-analysis that Conscientiousness is a universally valid predictor of job performance across all occupational categories, along with Agreeableness with similarly pervasive prediction (Barrick, M. R., & Mount, M. K. 1991)

https://psycnet.apa.org/doi/10.1111/j.1744-6570.1991.tb00688.x

https://psycnet.apa.org/record/1991-22928-001 

Summary of decades of research into selection methods finding cognitive ability among the strongest and most generalisable methods to predict future job performance (Schmidt, F. L., & Hunter, J. E. 1998) 

Summarised decades of cumulative selection research, proving general mental ability (g) or cognitive ability (or reasoning ability) is the single strongest overall predictor of job performance and learning velocity, along with stuctured interviews, and also including other strong but less generalisable and less convenient methods such as job sample testing (Schmidt, F. L., & Hunter, J. E. 1998) 

If you have read this far and are scrutinising this specific research entry, welcome friend and peer :) Before you smash that charlatan button, yes I omitted their 2016 update from Frank Schmidt with Oh and Shaffer, and the review from Sackett et al from 2021/2. I did this purposefully to avoid muddying the waters for this already large Resources entry. I am always down for a chat about it though, so make sure you head over to the contact page if you do too.

https://psycnet.apa.org/doi/10.1037/0033-2909.124.2.262

https://psycnet.apa.org/record/1998-10661-006 

Unified the CHC theory of human intelligence (essentially the periodic table of human intelligence) (McGrew, K. S. 1997)

Synthesised the extended Horn-Cattell Gf-Gc model with Carroll’s Three-Stratum Theory into a unified operational framework, formally creating what is now known as Cattell-Horn-Carroll (CHC) Theory, and informing modern cognitive ability testing in selection environments today (McGrew, K. S. 1997)

http://www.iapsych.com/chcbrief.htm

https://psycnet.apa.org/record/1997-97010-009

https://www.researchgate.net/publication/232605380_Analysis_of_the_major_intelligence_batteries_according_to_a_proposed_comprehensive_Gf-Gc_framework 

Tags:assessmentdevelopmentfeedbackperformancepersonalitypsychometricsselectiontools