Main Article Content

Abstract

Assessment of basic numeracy is essential for identifying students' foundational mathematical competence and for supporting instructional decisions. However, very high test scores may conceal limited measurement sensitivity when an instrument contains many low-difficulty items. This study examined an existing 25-item basic numeracy assessment by interpreting highly concentrated score distributions and proportion-correct values (p-values) as preliminary indicators of item difficulty. A descriptive quantitative approach was applied to data from 35 participants. Because the available dataset did not support formal reliability estimation, item discrimination indices, distractor analysis, or IRT/CDM modelling, the analysis was limited to total-score distribution, item-level proportion-correct values, and documentary review of item explanations. Results showed strong upper-end score concentration: scores ranged from 19 to 25 (M = 23.6, SD = 1.83), and 51.4% of participants achieved the maximum score. Items 11, 18, and 19 were answered correctly by all participants, whereas Items 24 and 10 produced the lowest correct-response rates but remained easy overall (85.7% and 88.6%). Item 17 was flagged for possible ambiguity in the answer explanation. The findings indicate that the instrument functions adequately as a baseline mastery check but has limited capacity to differentiate higher-performing students. The study contributes practical evidence for revising classroom-based numeracy instruments through more balanced item difficulty, clearer keying logic, and future validation using larger samples and stronger psychometric evidence.

Keywords

item analysis proportion-correct p-value ceiling effect item difficulty score distribution basic numeracy assessment educational measurement

Article Details

How to Cite
Fauziah, L., Rufi’i, R., & Sabariah, S. (2026). Beyond High Scores: Evaluating Item Difficulty and Measurement Sensitivity in Basic Numeracy Assessment. Jurnal Penelitian : Politeknik Penerbangan Surabaya, 20(1), 51–62. https://doi.org/10.46491/jp.v20i1.2547

References

  1. Feskens, R., Fox, J., & Zwitser, R. (2019). Differential item functioning in PISA due to mode effects. In M. von Davier, E. Gonzalez, I. Kirsch, & K. Yamamoto (Eds.), The role of international large-scale assessments: Perspectives from technology, economy, and educational research (pp. 231–247). Springer. https://doi.org/10.1007/978-3-030-18480-3_12
  2. Howard, S. J., Woodcock, S., Ehrich, J., & Bokosmaty, S. (2016). What are standardized literacy and numeracy tests testing? Evidence of the domain-general contributions to students' standardized educational test performance. British Journal of Educational Psychology, 87(1), 108–122. https://doi.org/10.1111/bjep.12138
  3. Le, H. V., & Nguyen, L. Q. (2024). A comparative study of critical reading abilities among students in Malaysia and Vietnam: Insights from PISA-based assessment. Research in Comparative and International Education, 19(2), 153–174. https://doi.org/10.1177/17454999241242994
  4. Lin, M., Bumgarner, E., & Chatterji, M. (2014). Understanding validity issues in international large-scale assessments. Quality Assurance in Education, 22(1), 31–41. https://doi.org/10.1108/qae-12-2013-0050
  5. McMillan, J. H. (2026). Classroom assessment validation: Proficiency claims and uses. Educational Measurement: Issues and Practice, 45(1), Article e70014. https://doi.org/10.1111/emip.70014
  6. Mor, E., & Karatoprak Erşen, R. (2023). Implications of current validity frameworks for classroom assessment. International Journal of Assessment Tools in Education, 10(Special Issue), 164–173. https://doi.org/10.21449/ijate.1368458
  7. Opesemowo, O. A. G., Opatunji, K. O., Babatimehin, T., & Opesemowo, T. R. (2026). Analysis of 2022 and 2023 Osun State basic education certificate examination mathematics items using item response theory: Implications for large scale assessment. Social Sciences & Humanities Open, 13, Article 102381. https://doi.org/10.1016/j.ssaho.2025.102381
  8. Rezigalla, A. A., Eleragi, A. M. E. S. A., Elhussein, A. B., Alfaifi, J., ALGhamdi, M. A., Al Ameer, A. Y., Yahia, A. I. O., Mohammed, O. A., & Adam, M. I. E. (2024). Item analysis: The impact of distractor efficiency on the difficulty index and discrimination power of multiple-choice items. BMC Medical Education, 24, Article 445. https://doi.org/10.1186/s12909-024-05433-y
  9. Rijmen, F., Jeon, M., von Davier, M., & Rabe-Hesketh, S. (2014). A third-order item response theory model for modeling the effects of domains and subdomains in large-scale educational assessment surveys. Journal of Educational and Behavioral Statistics, 39(4), 235–256. https://doi.org/10.3102/1076998614531045
  10. Woodcock, S., Howard, S. J., & Ehrich, J. (2020). A within-subject experiment of item format effects on early primary students' language, reading, and numeracy assessment results. School Psychology, 35(1), 80–87. https://doi.org/10.1037/spq0000340
  11. Wu, X., Li, N., Wu, R., & Liu, H. (2025). Cognitive analysis and path construction of Chinese students' mathematics cognitive process based on CDA. Scientific Reports, 15(1). https://doi.org/10.1038/s41598-025-89000-5
  12. Wu, X., Wu, R., Chang, H.-H., Kong, Q., & Zhang, Y. (2020). International comparative study on PISA mathematics achievement test based on cognitive diagnostic models. Frontiers in Psychology, 11. https://doi.org/10.3389/fpsyg.2020.02230