Psychology 214 Research Methods & Psychometrics Exam Notes (Stellenbosch University)

Psychology 214 at Stellenbosch University sits at the core of psychological science because it trains students to think like researchers rather than passive consumers of findings. Research methods and psychometrics together explain how knowledge in psychology is generated, tested, measured, and improved, from experimental design and sampling to reliability, validity, and test construction. Strong performance in this module depends on understanding both the logic of scientific inquiry and the practical mechanics of measurement.

1. The Logic of Psychological Research

Psychological research begins with a question, but a good question is never just a topic; it is a problem framed so that evidence can answer it. In Psychology 214, the central idea is that psychology is an empirical discipline: claims about behaviour, cognition, emotion, and personality must be supported by observable data. This makes the research process highly structured, because each stage from theory to conclusion must reduce ambiguity and limit bias.

1.1 Scientific thinking in psychology

Scientific thinking in psychology requires disciplined curiosity. A researcher does not simply ask, “Why do people behave this way?” and stop there. Instead, the researcher breaks the question into testable parts. For example, if one wants to know whether sleep deprivation reduces memory performance among Stellenbosch students, the idea must be translated into operational terms: What counts as sleep deprivation? How will memory be measured? Which students are included? Which alternative explanations must be controlled?

This emphasis on precision is essential because psychological phenomena are often invisible or indirectly observed. Unlike the height of a desk, which can be measured directly, constructs such as stress, intelligence, motivation, and anxiety require indicators. Research methods therefore exist to convert abstract ideas into measurable variables while still preserving meaning. The best studies are those where the operational definition is clear enough for replication and conceptually faithful enough to matter.

Scientific psychology also depends on systematic doubt. Findings are always provisional, because any result may be influenced by sampling error, instrument weakness, confounds, or statistical misinterpretation. Psychology 214 encourages a critical stance toward all claims, including published ones. A study with a significant result may still be poorly designed; a study with a non-significant result may still reveal an important effect that was underpowered or badly measured. The task is not merely to accept results but to evaluate how much confidence they deserve.

1.2 The research process from idea to conclusion

A typical research process in psychology follows a sequence:

  1. Identify a problem or phenomenon
  2. Review existing literature
  3. Formulate a research question
  4. Develop hypotheses
  5. Choose an appropriate design
  6. Select participants and sampling procedures
  7. Measure variables
  8. Collect data
  9. Analyse data
  10. Interpret findings
  11. Report conclusions and limitations

Each step shapes the quality of the next. If the literature review is weak, the research question may duplicate existing work or miss the theoretical contribution. If the sampling procedure is poor, the findings may not generalise. If the measurement is unreliable, the statistical analysis can become misleading even if the maths is correct.

The literature review is especially important because it situates a study in the broader scholarly conversation. It shows whether a variable has already been examined, whether findings have been contradictory, and what methodological gaps remain. In practical terms, a good research question often emerges from inconsistency: one study suggests a relationship, another does not, and the researcher asks why. Such gaps are fertile ground for psychological inquiry.

1.3 Hypotheses, variables, and operational definitions

A hypothesis is a specific, testable prediction about the relationship between variables. It is narrower than a theory and more concrete than a general idea. For example:

  • Students who sleep fewer than six hours before an exam will score lower on a memory test than students who sleep at least seven hours.
  • Higher test anxiety will be associated with lower exam performance.
  • Participants exposed to a stressful task will report higher state anxiety than participants in a neutral condition.

Every hypothesis depends on variables. A variable is any characteristic that can vary across people, groups, or time. Variables may be:

  • Independent variables (IVs): presumed causes, predictors, or manipulated conditions
  • Dependent variables (DVs): outcomes or responses
  • Control variables: factors held constant or statistically controlled
  • Confounding variables: unwanted influences that distort the causal interpretation

The crucial bridge between abstract theory and empirical research is the operational definition. This defines how a variable is measured or manipulated. “Anxiety” might be operationalised as a score on a validated self-report scale, physiological arousal, or behavioural avoidance. “Academic performance” might be operationalised as a percentage on a test, GPA, or score on a standardised task.

Operationalisation is not a trivial technicality. It determines what the study actually measures. If a researcher says they are studying “stress” but uses only one crude question, the study may capture annoyance rather than stress. Good researchers match the operational definition to the theoretical construct as closely as possible.

1.4 Induction, deduction, and theory testing

Research in psychology uses both deductive and inductive reasoning. Deduction starts with a theory and tests what should happen if the theory is correct. Induction begins with observation and builds broader generalisations from patterns in data.

In practice, psychology often moves back and forth between the two:

  • A theory predicts that test anxiety impairs working memory.
  • A study tests this prediction with students under exam-like conditions.
  • If results support the theory, confidence in the theory increases.
  • If results contradict it, the theory may be revised or narrowed.

This iterative process is how psychological science advances. No single study proves a theory once and for all. Instead, many studies accumulate evidence, refine constructs, and identify boundary conditions. A theory may work well for one age group, one cultural context, or one setting but fail in another. Psychologically meaningful generalisations therefore require repeated testing across contexts.

1.5 Common sources of bias in research thinking

Psychology 214 places strong emphasis on recognising bias because human beings are not neutral observers. Common biases include:

  • Confirmation bias: favouring evidence that supports pre-existing beliefs
  • Availability bias: overestimating the importance of vivid or recent examples
  • Demand characteristics: participants altering behaviour based on perceived expectations
  • Experimenter expectancy effects: researchers unintentionally influencing outcomes
  • Publication bias: studies with significant findings are more likely to be published

These biases matter because they can make an idea look stronger than it really is. For example, a researcher expecting caffeine to improve attention might unconsciously give more encouragement to the caffeine group or interpret ambiguous responses more positively. Similarly, participants may try to behave in ways they think the researcher wants. Good design and careful measurement are tools for reducing these threats.

A strong research mindset is therefore both creative and sceptical. It is creative because it asks original questions and proposes testable explanations. It is sceptical because it knows that evidence can be distorted by poor design, weak measurement, and human bias. Psychology 214 essentially trains students to combine these habits into a disciplined form of inquiry.

2. Research Designs, Sampling, and Threats to Validity

Research design determines what kinds of conclusions a study can support. Different designs answer different questions, and no design is universally best. A good method is one that matches the research question, the available resources, ethical constraints, and the level of inference needed. In psychology, design choice is often the difference between describing a pattern, explaining a relationship, or demonstrating causation.

2.1 Descriptive, correlational, and experimental designs

The three broad design families are descriptive, correlational, and experimental.

Descriptive designs

Descriptive studies aim to observe and summarise behaviour without manipulating variables. Examples include case studies, naturalistic observation, surveys, and archival analyses. These designs are useful when the phenomenon is not yet well understood or when manipulation would be unethical.

A descriptive study might examine how first-year students at Stellenbosch University describe their coping strategies during exam season. Such a study can reveal patterns, frequencies, and lived experiences, but it cannot by itself show that one factor causes another.

Correlational designs

Correlational studies examine whether two or more variables move together. A positive correlation means that as one variable increases, the other tends to increase; a negative correlation means that as one increases, the other tends to decrease. Correlation is valuable for prediction and theory development, but it does not imply causation.

For example, if test anxiety and exam failure are correlated, several explanations are possible: anxiety may reduce performance, poor performance may increase anxiety, or both may be influenced by a third variable such as low preparation or low self-efficacy. Psychology 214 repeatedly stresses this point because correlation is one of the most commonly misunderstood concepts in social science.

Experimental designs

Experimental studies involve manipulation of an independent variable and control of extraneous variables, usually through random assignment. This is the strongest design for causal inference because it helps isolate the effect of the manipulated factor.

Consider a study comparing memory performance after three hours of sleep versus seven hours of sleep. If participants are randomly assigned to sleep conditions and other influences are controlled, differences in memory can be more confidently attributed to sleep duration. However, experiments are not always easy to implement in real settings, and ethical limits are significant.

2.2 Variables, control, and confounding

A central task in design is distinguishing a real effect from alternative explanations. This is where control becomes essential. Control can be achieved through:

  • Random assignment
  • Standardised procedures
  • Matching participants
  • Holding variables constant
  • Statistical control
  • Blinding

A confound occurs when an extraneous variable varies systematically with the independent variable and offers a competing explanation for the results. For instance, if one group in a study on teaching methods is taught by a highly experienced lecturer and the other by a new lecturer, lecturer experience becomes a confound. Any difference in student performance might reflect the teacher rather than the teaching method.

The stronger the control, the stronger the internal validity. But there is often a trade-off: highly controlled studies may feel artificial and less representative of everyday behaviour. This trade-off is a persistent issue in psychology, especially when the subject matter is social, contextual, or emotionally charged.

2.3 Sampling and generalisability

Sampling determines whether findings can be extended beyond the individuals who actually participated. In theory, researchers want the sample to represent the target population. In practice, samples are often constrained by convenience, access, and cost.

Common sampling methods include:

  • Simple random sampling: every member of the population has an equal chance of selection
  • Stratified sampling: the population is divided into subgroups, and participants are sampled from each subgroup
  • Systematic sampling: selecting every nth member from a list
  • Convenience sampling: selecting those who are easiest to access
  • Purposive sampling: selecting participants with particular characteristics
  • Snowball sampling: participants recruit additional participants from their networks

Each method has strengths and weaknesses. Random sampling is ideal for representativeness but often unrealistic. Convenience sampling is common in university research because it is feasible, but it limits generalisability. For example, using only first-year psychology students as participants may tell us a great deal about that group but very little about older adults, working professionals, or people in different cultural settings.

A useful distinction in Psychology 214 is between the sample and the population. The sample is the actual group studied; the population is the broader group to which the researcher wants to generalise. Good research makes the connection between the two explicit.

2.4 Internal, external, construct, and statistical validity

Validity is not a single concept but a family of related concerns.

Internal validity

Internal validity is the extent to which a study supports a causal conclusion. It is threatened by confounds, selection bias, maturation, history, testing effects, instrumentation changes, and attrition.

External validity

External validity concerns whether findings generalise beyond the study sample, setting, or procedure. A lab study may have high internal validity but low external validity if it is too artificial.

Construct validity

Construct validity concerns whether the study truly measures or manipulates the theoretical construct of interest. If a test is meant to measure depression but actually measures fatigue and loneliness more strongly, construct validity is weak.

Statistical conclusion validity

Statistical conclusion validity refers to whether the statistical inference is appropriate. Problems include low power, assumption violations, unreliable measures, and poor statistical choice.

These four forms of validity are often intertwined. A study with weak construct validity may produce a misleading effect even if its statistics are sound. A study with poor internal validity may produce a strong-looking result that cannot support causation. Psychology 214 therefore encourages students to evaluate studies holistically, not just by whether a p-value is less than 0.05.

2.5 Ethics in psychological research

Ethics is inseparable from methodology. Psychological research often involves human vulnerability, privacy, deception, sensitive data, or power asymmetries. Ethical principles commonly include:

  • Informed consent
  • Voluntary participation
  • Right to withdraw
  • Protection from harm
  • Confidentiality and anonymity
  • Debriefing after deception
  • Fair participant selection

Ethical decisions are not merely bureaucratic. They protect dignity and improve data quality. Participants who feel respected are more likely to respond honestly. Conversely, unethical procedures can produce distress, dropout, or invalid responses. In contexts involving students, staff, or marginalised groups, ethical sensitivity is especially important because participation may be influenced by implicit pressure or unequal power relationships.

A frequent exam point is that ethical constraints sometimes shape design choice. For example, one cannot randomly assign people to severe trauma. This means many important psychological questions must be studied through naturalistic, correlational, or retrospective designs rather than experiments. Methodological realism is therefore a response to ethical reality.

3. Measurement, Scales, and Psychometrics

Psychometrics is the science of psychological measurement. It asks whether instruments actually measure what they claim to measure and how consistently they do so. In Psychology 214, psychometrics is not an isolated topic; it is the measurement foundation for all research methods. If the measurement tool is weak, even the most elegant research design will produce shaky conclusions.

3.1 Why measurement matters

Psychological constructs are often latent, meaning they cannot be observed directly. We do not see intelligence, anxiety, or self-esteem in a physical sense; we infer them from patterns of responses, behaviour, or performance. Because inference is involved, measurement error is always present to some degree.

Measurement quality affects every stage of research:

  • It influences whether variables are distinguishable.
  • It affects statistical power.
  • It determines whether group differences are meaningful.
  • It shapes the credibility of conclusions.
  • It affects whether a scale can be used in other contexts.

A study on depression that uses a poor measure may fail even if the hypothesis is correct. If the instrument misses core symptoms, uses ambiguous language, or is culturally inappropriate, the results may reflect measurement failure rather than psychological reality.

3.2 Levels of measurement

Understanding measurement scales is essential because the level of measurement determines what analyses and interpretations are appropriate.

Level of measurement Meaning Example in psychology Appropriate comparisons
Nominal Categories with no inherent order Gender categories, diagnosis groups Frequencies, modes, chi-square
Ordinal Ordered categories without equal intervals Likert-type responses, rank order Median, non-parametric tests
Interval Equal intervals, no true zero Some standardised scores Means, standard deviations, correlation
Ratio Equal intervals with a true zero Reaction time, number correct Full arithmetic operations

A common exam issue is the treatment of Likert scales. A single Likert item is technically ordinal, because the distance between “agree” and “strongly agree” is not guaranteed to be equal to the distance between other categories. However, summed Likert scales with several items are often treated as approximately interval in practice, especially when the scale has good psychometric properties and the distribution is acceptable. This is a pragmatic research convention, not a mathematical truth.

3.3 Reliability: consistency of measurement

Reliability refers to the consistency, stability, or dependability of a measurement instrument. A scale can be reliable without being valid, but a scale cannot be valid if it is highly inconsistent. Reliability can be examined in several ways:

  • Test-retest reliability: stability over time
  • Internal consistency: how well items on a scale relate to one another
  • Inter-rater reliability: agreement between observers or coders
  • Parallel-forms reliability: consistency across equivalent versions of a test

Test-retest reliability

If a personality inventory is administered to the same participants two weeks apart, similar scores suggest test-retest reliability. This is important for traits that should be relatively stable. Large changes may indicate instability or poor measurement.

Internal consistency

Internal consistency asks whether items intended to measure the same construct produce similar responses. A scale for social anxiety should contain items that relate coherently to social fear, avoidance, or discomfort. If the items are too diverse or poorly worded, consistency drops. Cronbach’s alpha is commonly used as an index of internal consistency, though it must be interpreted with caution because very high alpha values may reflect item redundancy rather than genuine breadth.

Inter-rater reliability

When human judgement is involved, reliability depends on agreement between raters. For example, if two observers rate non-verbal aggression in classroom interactions, their coding must be sufficiently aligned. Training, clear coding manuals, and pilot work are essential to improve agreement.

3.4 Validity: measuring the right thing

While reliability asks whether a measure is consistent, validity asks whether it actually measures the intended construct. The main forms include:

  • Face validity: does the measure appear appropriate on the surface?
  • Content validity: does the instrument cover the full domain of the construct?
  • Criterion validity: does it relate to an external criterion?
  • Construct validity: does it behave as theory predicts?
  • Convergent validity: does it correlate with related measures?
  • Discriminant validity: does it not correlate too strongly with unrelated measures?

Construct validity is the most comprehensive and the most important in advanced psychological measurement. It develops through an accumulation of evidence, not a single statistic. For example, a new anxiety scale may show convergent validity if it correlates with other anxiety measures, discriminant validity if it does not overlap excessively with unrelated traits such as openness, and criterion validity if it predicts avoidance behaviour in stressful situations.

Validity is also contextual. A scale that works well in one language or culture may not work equally well elsewhere. Translation is not enough; meaning must be preserved. This is particularly relevant in South African research, where multilingual and multicultural contexts make measurement equivalence a major concern.

3.5 Psychometric item quality and scale construction

Good psychometrics depends on careful item writing. Items should be:

  • Clear and unambiguous
  • Single-barrelled
  • Appropriate to the respondent’s reading level
  • Free from double negatives
  • Balanced in tone where necessary
  • Focused on one construct at a time

Poor item examples include:

  • “I am often anxious and unhappy.”
    This combines two constructs.
  • “I never do not feel stressed.”
    This double negative is confusing.
  • “I am a good person in all situations.”
    This is vague and socially desirable.

Scale construction usually follows a systematic process:

  1. Define the construct theoretically.
  2. Generate a large pool of items.
  3. Review items for content coverage.
  4. Pilot the scale.
  5. Analyse item performance.
  6. Remove weak or ambiguous items.
  7. Test reliability and validity again.
  8. Cross-validate in a new sample.

Item analysis may look at item-total correlations, discrimination, difficulty, and response patterns. A good item distinguishes participants who are high on the construct from those who are low on it. If an item does not contribute meaningfully to the scale, it may need revision or removal.

3.6 Standardisation and norms

A test becomes especially useful when it is standardised. Standardisation means that administration, scoring, and interpretation procedures are uniform. Norms are the reference values used to interpret an individual’s score relative to a group. For example, a score may be interpreted relative to age-based norms, percentile ranks, or standard scores.

Norms matter because a raw score is often meaningless on its own. A score of 28 on a stress inventory can only be interpreted if one knows what 28 means relative to a reference group and the instrument’s scoring system. In psychometrics, interpretation is always relational.

3.7 Measurement error and why no score is perfect

Every observed score contains some degree of true score and some degree of error. Error can arise from fatigue, misunderstanding, distraction, guessing, scoring mistakes, or temporary emotional states. Measurement error reduces the observed association between variables and can make real relationships harder to detect.

This is why strong psychometric instruments are indispensable in research methods. They increase the likelihood that observed differences reflect actual differences rather than noise. For Psychology 214, this means that psychometric reasoning is not separate from research design; it is the condition that makes design meaningful.

4. Statistics, Data Analysis, and Interpretation

Statistics in Psychology 214 are not just mathematical procedures; they are tools for deciding what the data can legitimately say. Good statistical practice begins long before any formula is applied. It starts with understanding the research question, the type of variables involved, the design structure, and the assumptions behind the analysis. A statistical test is only as good as the logic that supports its use.

4.1 Descriptive statistics

Descriptive statistics summarise data in a manageable form. Common descriptive measures include:

  • Mean: the arithmetic average
  • Median: the middle value
  • Mode: the most frequent value
  • Range: highest minus lowest score
  • Variance: the average squared deviation from the mean
  • Standard deviation: the square root of variance, showing spread
  • Percentiles: relative standing in a distribution

Descriptive statistics are important because they reveal the shape and spread of the data. A mean can be misleading if the distribution is skewed. For example, if most students score around 60 but a few score very low, the mean may not fully represent the typical performance. In such cases, the median may better reflect the centre of the distribution.

Graphical displays are equally important:

  • Histograms show distributions.
  • Boxplots highlight medians, spread, and outliers.
  • Scatterplots show relationships between variables.
  • Bar charts compare categories.

Before any inferential analysis, researchers should inspect the data visually. Outliers, impossible values, skewness, and data entry errors often become obvious only at this stage.

4.2 Inferential statistics and sampling logic

Inferential statistics allow researchers to make statements beyond the sample. They rest on the logic that if a sample is representative and the analysis is appropriate, the observed pattern can be used to estimate population tendencies. This is why random sampling and random assignment are so often emphasised: they support inference in different ways.

The two broad questions of inferential statistics are:

  1. Is there evidence of an effect or relationship?
  2. How large and practically important is it?

This distinction matters because statistical significance is not the same as substantive importance. A tiny effect can be statistically significant in a very large sample, while a practically important effect may fail to reach significance in a small sample.

4.3 Hypothesis testing, p-values, and Type I/II errors

Hypothesis testing typically begins with a null hypothesis stating that there is no effect or no difference. The researcher then asks whether the observed data are unusual enough under the null hypothesis to justify rejecting it.

The p-value is the probability of obtaining results at least as extreme as the observed data, assuming the null hypothesis is true. A small p-value suggests that the observed result would be unlikely if there were truly no effect.

However, the p-value does not tell us:

  • The probability that the hypothesis is true
  • The size of the effect
  • The practical importance of the finding
  • Whether the design was good

Common error types include:

  • Type I error: rejecting a true null hypothesis; a false positive
  • Type II error: failing to reject a false null hypothesis; a false negative

The balance between these errors depends partly on the chosen significance level and the statistical power of the study. Lowering the alpha level reduces Type I error but increases the risk of Type II error unless sample size or power improves.

4.4 Statistical power and sample size

Power is the probability of detecting an effect if it truly exists. Low power is a major problem because it makes studies less likely to detect meaningful effects and can inflate the proportion of false positives among significant results. Power depends on:

  • Sample size
  • Effect size
  • Variability in the data
  • Alpha level
  • Measurement reliability

A study with excellent design but too few participants may still be inconclusive. This is particularly relevant in psychology, where access to participants may be limited and effects are often moderate or small. Small samples also produce unstable estimates, making replication more difficult.

Researchers should therefore think about sample size early, not as an afterthought. A well-powered study is ethically stronger too, because it avoids exposing participants to procedures that are unlikely to answer the question adequately.

4.5 Effect sizes and practical significance

Effect sizes quantify the magnitude of a relationship or difference. Unlike p-values, effect sizes help answer the question: how big is the effect? Common examples include:

  • Cohen’s d for mean differences
  • Correlation coefficients for linear relationships
  • Eta-squared or partial eta-squared for explained variance in ANOVA contexts
  • Odds ratios for categorical outcomes

Effect sizes are crucial because a statistically significant effect may still be too small to matter in practice. For example, a teaching intervention might raise average test scores by half a point on a 100-point exam. If the sample is huge, that difference may be statistically significant but educationally trivial. By contrast, a larger effect with a modest sample might fail to reach significance yet still deserve attention.

A mature interpretation combines significance, effect size, confidence intervals, and research design quality.

4.6 Common statistical tests and when to use them

The choice of test depends on the research question and the type of data. Common tests include:

Research question Typical test Main use
Compare two independent group means t-test Difference between two groups
Compare more than two group means ANOVA Differences among several groups
Examine relationship between two continuous variables Correlation Strength and direction of association
Predict one variable from another or several others Regression Prediction and modelling
Compare categorical frequencies Chi-square Association between categories
Examine within-subject changes across time Paired t-test / repeated-measures ANOVA Pre-post or repeated measures

These tests are not interchangeable. Using the wrong test can distort conclusions. For example, a chi-square test is appropriate for categories, not for mean differences between continuous scores. Likewise, choosing a t-test when there are more than two groups can inflate Type I error.

4.7 Assumptions and interpretation

Most inferential tests rest on assumptions such as:

  • Independence of observations
  • Approximate normality
  • Homogeneity of variance
  • Linearity for correlations and regression
  • Adequate expected frequencies for chi-square

Assumptions are not mere technical details. They determine whether a test is trustworthy. If the assumptions are seriously violated, the result may be misleading, even if the software returns a p-value.

Interpretation should also be cautious about causality. Correlation does not imply causation because of:

  • Reverse causality
  • Third-variable explanations
  • Measurement overlap
  • Spurious relationships

For example, if social media use correlates with stress, one cannot conclude immediately that social media causes stress. It is equally plausible that stressed students use social media differently, or that lack of sleep influences both.

4.8 Reading results critically

A strong Psychology 214 student reads results in a layered way:

  1. What was the research question?
  2. Was the design suitable?
  3. Were the measures reliable and valid?
  4. Was the sample appropriate?
  5. Were the statistical tests correct?
  6. Was the effect size meaningful?
  7. What are the limitations?
  8. Do the conclusions go beyond the data?

This disciplined reading prevents overclaiming. A study can produce a neat table of results and still be weak in its logic. The best interpretation combines numbers with methodological judgement.

5. Revision Themes, Common Exam Questions, and Application Strategies

Exam success in Psychology 214 depends on more than memorising definitions. Students must be able to compare concepts, apply them to scenarios, and explain why a method or statistic is appropriate. The module rewards conceptual understanding, especially where research design and psychometrics intersect. A good answer often identifies not just the correct term, but also the reason it fits the situation better than alternatives.

5.1 High-yield concepts to master

The most examinable ideas usually include:

  • Difference between correlation and causation
  • Difference between reliability and validity
  • Difference between internal and external validity
  • Difference between sample and population
  • Difference between independent and dependent variables
  • Difference between nominal, ordinal, interval, and ratio data
  • Difference between Type I and Type II errors
  • Relationship between effect size, sample size, and power
  • How operational definitions shape conclusions
  • Why random assignment improves causal inference
  • Why random sampling improves generalisability

These are not isolated facts. They form a conceptual web. For example, a weak measure reduces reliability, which undermines validity, which weakens construct interpretation, which in turn affects the value of any statistical test conducted on the data.

5.2 Common scenario types and how to think through them

Psychology examinations often present a short scenario and ask students to identify the design, variables, validity threats, or suitable statistics. The best strategy is to slow down and classify the essential features.

Scenario type 1: group comparison

If a study compares two teaching methods and measures exam scores, the key questions are:

  • Is the independent variable manipulated or naturally occurring?
  • Are participants randomly assigned?
  • Is the dependent variable continuous?
  • Are there more than two groups?

If there are two independent groups and a continuous outcome, a t-test may be appropriate. If there are more than two groups, ANOVA may be better.

Scenario type 2: relationship between variables

If a study examines whether study hours relate to GPA, the question is whether both are continuous and whether the goal is association or prediction. Correlation is suitable for association; regression is useful for prediction and when multiple predictors are involved.

Scenario type 3: psychometric scale evaluation

If a new scale is introduced, ask:

  • Is the scale internally consistent?
  • Does it show evidence of validity?
  • Is the sample appropriate for the target population?
  • Are the items worded clearly?
  • Is the scale culturally and linguistically suitable?

A scale may look professional yet still be psychometrically weak if these issues are ignored.

5.3 Revision table: quick comparison of core ideas

Concept pair Key difference Why it matters
Reliability vs validity Consistency vs accuracy A consistent measure can still be wrong
Internal vs external validity Causal confidence vs generalisability Good science needs both, but trade-offs are common
Random sampling vs random assignment Generalise to population vs support causal inference They solve different problems
P-value vs effect size Statistical evidence vs magnitude Significance alone is not enough
Correlation vs causation Association vs causal explanation Prevents overclaiming
Nominal vs ordinal data Categories vs ranked categories Determines suitable analysis

This table is especially helpful for exam revision because many questions ask students to discriminate between similar-sounding concepts.

5.4 How to answer longer exam questions

A strong long-form answer usually follows a structure:

  1. Define the concept clearly
  2. Explain its purpose
  3. Describe how it works in research
  4. Give an example
  5. Mention a limitation or related concept
  6. Conclude with why it matters

For example, if asked about reliability, one should not stop at “reliability means consistency.” A fuller answer would explain why it matters for measurement, mention test-retest or internal consistency, note that reliability is necessary but not sufficient for validity, and apply the concept to a psychological test or questionnaire.

Similarly, for validity, a strong answer should distinguish forms of validity and show that validity is evidence-based, not a single checklist item. If asked about sampling, one should distinguish representativeness from practicality and note how sampling decisions affect external validity.

5.5 Common mistakes to avoid

Students often lose marks for the same recurring reasons:

  • Using correlation language when the design is actually experimental
  • Confusing random sampling with random assignment
  • Treating p < 0.05 as proof of importance
  • Ignoring the role of measurement error
  • Assuming a reliable test must also be valid
  • Describing a variable as if it is automatically continuous when it may be ordinal
  • Failing to mention limitations in a design
  • Calling every self-report questionnaire a “test” without discussing its psychometric quality

Another common mistake is to overload an answer with technical terms but little explanation. Examiners usually reward clarity and logical flow more than jargon. A concise, correct explanation beats a long string of undeveloped terms.

5.6 Integrating research methods and psychometrics

The real strength of Psychology 214 is that it teaches research methods and psychometrics as mutually dependent. A study is only as good as its design, and a design is only as good as its measurement. This is why the module should not be studied as separate compartments. Sampling affects inference, design affects causality, and psychometrics affects whether the variables themselves are trustworthy.

A useful way to remember the integration is:

  • Research methods answer: How do we produce evidence?
  • Psychometrics answers: How do we measure psychological constructs well?

Together they determine whether psychological knowledge is credible. A good study asks a meaningful question, uses an appropriate design, samples wisely, measures accurately, analyses correctly, and interprets cautiously. That chain is the backbone of psychological science at Stellenbosch University and the reason this module is so central to later work in psychology.

5.7 Final revision priorities

Before an exam, revision should focus on mastering both definitions and application. The most productive study routine usually includes:

  • Rewriting definitions in your own words
  • Practising scenario-based classification
  • Comparing similar concepts side by side
  • Drawing simple diagrams of designs and variable relationships
  • Memorising the logic, not just the labels, of statistical tests
  • Reviewing psychometric terms with examples of good and poor instruments
  • Testing yourself on why a particular design or measure is appropriate

In the end, success in Psychology 214 comes from understanding that psychological science is a craft of careful judgement. It asks students to think critically, measure responsibly, analyse honestly, and conclude modestly. Those habits are what make research methods and psychometrics not only examinable content, but the foundation of credible psychology.

Select the fields to be shown. Others will be hidden. Drag and drop to rearrange the order.
  • Image
  • SKU
  • Rating
  • Price
  • Stock
  • Availability
  • Add to cart
  • Description
  • Content
  • Weight
  • Dimensions
  • Additional information
Click outside to hide the comparison bar
Compare