Read this report with me
A hands-on lesson on the ready-made sample studies: we open the results of a realistic study and read each figure in the order an examiner would, learning what belongs in a thesis and what doesn't.
Why this lesson
Real data gives whatever it gives, and many tests come back empty, so it teaches you little about what the tests do. The sample studies are built on purpose so every test has something worth reading: a reliable construct and a failing one, a reverse-worded item, and group differences only post-hoc comparisons can reveal. The data comes from a fixed-seed generator, so the same figures appear on every device and you can check what you see against this page.
Note The data is for teaching, not real, and each study's description says so.
The four studies
| Code | Study | Responses | Why it is here |
|---|---|---|---|
SHOWCASE-MAIN |
Academic engagement and achievement | 300 | Every test that needs a real sample |
SHOWCASE-PILOT |
Pilot study: pre-test | 16 | Exact tests for small samples |
SHOWCASE-POST |
Pilot study: post-test | 16 | The paired t-test |
SHOWCASE-BROKEN |
A deliberately broken study | 24 | When Pulseform refuses to give a number, and why |
If the studies are in your account, open Surveys, then View Responses & Results for
SHOWCASE-MAIN, then the Academic & Reliability (Cronbach) tab. If they aren't, read the figures here and
follow the same steps with your own data.
Step 1: Reliability first
| Construct | Alpha | Verdict |
|---|---|---|
| Behavioural engagement | 0.946 | Excellent |
| Cognitive engagement | 0.946 | Excellent |
| Emotional engagement | 0.948 | Excellent |
| Supporting factors | 0.482 | Rejected |
Supporting factors is built to fail: alpha is below the accepted 0.70, and even below 0.60. Open the items table and look at the corrected item-total correlation: it is low for every item. This is what an instrument that doesn't measure one idea looks like.
Warning Don't run any hypothesis test on a construct whose reliability is rejected. Report the weakness in your thesis, and reword its items for your next study.
Step 2: The reverse-worded item
The item "I skip lectures without a good reason" is negatively worded and marked in the editor with This item is reverse scored, which is why its correlation with its construct is positive.
Try it: Open the question in the editor, clear This item is reverse scored, and go back to the academic tab. Behavioural engagement's alpha drops from 0.946 to below 0.5, though not a single answer changed. Then tick it again. This shows why every reverse-worded item must be marked before you read any result.
Step 3: Is the data normal?
Skewness and kurtosis say the engagement constructs and the GPA are normal, yet Shapiro-Wilk rejects all of them. That isn't a mistake: with 300 respondents, Shapiro-Wilk rejects any departure, however small.
- Rely on skewness and kurtosis for large samples, and report Shapiro-Wilk with this caveat if your supervisor asks for it.
- Library hours really is skewed (+0.87), so the correlation matrix recommends Spearman. Remove it from the matrix and the recommendation switches to Pearson.
Step 4: Differences between groups
By gender (independent-samples t-test):
t(298) = 3.54, p < .001, d = 0.41, 95% CI [0.15, 0.52]
- The confidence interval doesn't contain zero: the same statement as
p < .001, in the scale's units. d = 0.41is a small-to-medium effect: the difference is significant but not large. That is the difference between significance and size.
By major: F(2, 297) = 1.24, p = .290. No significant difference, so no post-hoc comparisons are
shown. This is a finding to report as it is, not a "failed" result.
Step 5: Post-hoc comparisons tell the story
By study year: F(3, 296) = 10.60, p < .001, η² = 0.097.
ANOVA alone says "there are differences", which is half a finding. The post-hoc comparisons (Tukey and Scheffé) name the pairs: years one and two are alike, years three and four are alike, and the gap is between the two groups. Notice too that Tukey is always closer to significance than Scheffé, because it is the more powerful test.
Step 6: Multiple regression
F(3, 296) = 345.31, p < .001, R² = .778
| Predictor | β | p |
|---|---|---|
| Behavioural engagement | .613 | < .001 |
| Cognitive engagement | .384 | < .001 |
| Emotional engagement | .032 | .341 |
- Emotional engagement is not significant, and its confidence interval contains zero.
- VIF is between 1.45 and 1.50, so the ranking of the predictors is stable and quotable. Above 5, "which is strongest?" would have no stable answer.
Step 7: Factor analysis confirms what alpha said
A notice appears: "You declared 4 constructs; the data contains 3." That is right: the three engagement constructs form their own factors, while supporting factors forms none, as alpha (0.482) already said.
Tip This is a finding about your instrument, not an analysis error. You might write: "Factor analysis extracted three factors; the supporting-factors items did not load on a separate factor, consistent with their low reliability (α = 0.482)."
Step 8: Small samples and pre/post
- In
SHOWCASE-PILOT(8 per group), Pulseform computes the exact p-value (0.0014) by counting every possible split, instead of the approximation. This is SPSS's "Exact Sig.", and it is the one to report. - Comparing
SHOWCASE-PILOTwithSHOWCASE-POST, Pulseform pairs each person with themselves by respondent code:t(15) = 2.40, p = .030, d = 0.60. The means differ by only 0.42; treated as two independent samples, no significant difference would show. See Pre/post surveys.
Step 9: When Pulseform refuses
Open SHOWCASE-BROKEN. Every flaw in it is deliberate:
| What is in the data | What Pulseform says |
|---|---|
| Everyone gave the same answer on a construct | Alpha unacceptable, and correlation not computable, not "zero" |
| A one-item construct | Alpha not applicable |
| Two identical items | A perfect correlation with no confidence interval; regression and factor analysis refused |
| 23 men and one woman | Genders can't be compared |
| 6 majors among 24 people | The cross-tab is too sparse for chi-square |
A refusal is information, not a fault. A misleading number is worse than none: "r = 0" would have entered your thesis as a finding the data doesn't support. And note that a refusal describes the first obstacle the calculation hit, so narrow your selection to see the rest.
What's next
- Follow the same steps with your own data in Academic surveys.
- Look up any term in the glossary.