Skip to main navigation menu Skip to main content Skip to site footer

Articles

Vol. 28 No. 3 (2025): Second language pronunciation: Current issues and developments

Dictation automatic speech recognition for second language pronunciation assessment: Focus on age-related bias

DOI
https://doi.org/10.37213/cjal.2025.35706
Submitted
September 11, 2025
Published
2025-12-15

Abstract

Dictation automatic speech recognition (ASR) technology offers significant advantages in second language (L2) pronunciation assessment, a claim supported by research showing strong correlations between ASR and human-rater scores. However, ASR systems may introduce biases linked to age-related physiological changes that affect performance. Despite studies on age bias, none address dictation systems in assessment contexts. This study investigates age bias in Google Voice Typing (GVT), Microsoft-Transcribe (MS-T), and Apple Siri for English L2 pronunciation assessment. We analyzed 1,000 university test responses spanning five L1 groups (French, Spanish, Persian, Arabic, Chinese) and three age groups (<30, 30–44, >44). Recordings were processed through GVT, MS-T, and Siri to generate word accuracy scores. Regression analyses revealed the target ASR systems favored test takers ≤29 years (p =.001). These findings highlight dictation ASR-related age biases and underscore the need for fairness in automated assessments.

References

  1. Arora, V., Lahiri, A., & Reetz, H. (2018). Phonological feature-based speech recognition system for pronunciation training in non-native language learning. The Journal of the Acoustical Society of America, 143(1), 98–108. https://doi.org/10.1121/1.5017834 DOI: https://doi.org/10.1121/1.5017834
  2. Babaeian, A. (2023). Pronunciation assessment: Traditional vs modern modes. Journal of Education for Sustainable Innovation, 1(1), 61–68. https://doi.org/10.56916/jesi.v1i1.530 DOI: https://doi.org/10.56916/jesi.v1i1.530
  3. Badwan, K. (2021). Language and the sociolinguistic market. In Language in a globalised world: Social justice perspectives on mobility and contact. Springer International Publishing. https://doi.org/10.1007/978-3-030-77087-7_3 DOI: https://doi.org/10.1007/978-3-030-77087-7_3
  4. Bajorek, J. (2019). Voice recognition still has significant race and gender biases. Harvard Business Review, 10, 1–4.
  5. Bennett, R. E., Braswell, J., Oranje, A., Sandene, B., Kaplan, B., & Yan, F. (2008). Does it matter if I take my mathematics test on computer? A second empirical study of mode effects in NAEP. Journal of Technology, Learning, and Assessment, 6(9), 1–39.
  6. Bernstein, J., Van Moere, A., & Cheng, J. (2010). Validating automated speaking tests. Language Testing, 27(3), 355–377. https://doi.org/10.1177/0265532210364404 DOI: https://doi.org/10.1177/0265532210364404
  7. Bialystok, E., Craik, F., & Luk, G. (2008). Cognitive control and lexical access in younger and older bilinguals. Journal of Experimental Psychology: Learning, Memory, and Cognition, 34(4), 859–873. https://doi.org/10.1037/0278-7393.34.4.859 DOI: https://doi.org/10.1037/0278-7393.34.4.859
  8. Birdsong, D. (2006). Age and second language acquisition and processing: A selective overview. Language Learning, 56(S1), 9–49. https://doi.org/10.1111/j.1467-9922.2006.00353.x DOI: https://doi.org/10.1111/j.1467-9922.2006.00353.x
  9. Blake, R. (2016). Technology and the four skills. Language Learning & Technology, 12(3), 129–142. DOI: https://doi.org/10.64152/10125/44465
  10. Caldwell-Harris, C. L., & MacWhinney, B. (2023). Age effects in second language acquisition: Expanding the emergentist account. Brain and Language, 241. https://doi.org/10.1016/j.bandl.2023.105269 DOI: https://doi.org/10.1016/j.bandl.2023.105269
  11. Cámara-Arenas, E., Tejedor-García, C., Tomas-Vásquez, C. J., & Escudero-Mancebo, D. (2023). Automatic pronunciation assessment vs. Automatic speech recognition: A study of conflicting conditions for L2-English. Language Learning & Technology, 27(1), 1–19. DOI: https://doi.org/10.64152/10125/73512
  12. Chapelle, C. A., & Lee, H. (2021). Conceptions of validity. In G. Fulcher, & L. Harding, The Routledge handbook of language testing (2nd ed., pp. 17–31). Routledge. https://doi.org/10.4324/9781003220756-3 DOI: https://doi.org/10.4324/9781003220756-3
  13. Chapelle, C. A., & Voss, E. (2016). 20 Years of technology and language assessment in language learning & technology. Language Learning & Technology, 20(2), 116–128. DOI: https://doi.org/10.64152/10125/44464
  14. Chun, D., Kern, R., & Smith, B. (2016). Technology in language use, language teaching, and language learning. The Modern Language Journal, 100(S1), 64–80. https://doi.org/10.1111/modl.12302 DOI: https://doi.org/10.1111/modl.12302
  15. Coulmas, F. (2005). Communicating across generations: Age as a factor of linguistic choice. In Sociolinguistics: The study of speakers’ choices (1st ed., pp. 52–67). Cambridge University Press. https://doi.org/10.1017/CBO9780511815522 DOI: https://doi.org/10.1017/CBO9780511815522.004
  16. Cox, T. L., & Davies, R. S. (2012). Using automatic speech recognition technology with elicited oral response testing. CALICO Journal, 29(4), 601–618. https://doi.org/10.11139/cj.29.4.601-618 DOI: https://doi.org/10.11139/cj.29.4.601-618
  17. Cucchiarini, C., Strik, H., & Boves, L. (1997). Automatic evaluation of Dutch pronunciation by using speech recognition technology. In Proceedings of Eurospeech 1997 (pp. 61–68). https://doi.org/10.1109/ASRU.1997.659144 DOI: https://doi.org/10.1109/ASRU.1997.659144
  18. Derwing, T. M., & Munro, M. J. (1997). Accent, intelligibility, and comprehensibility: Evidence from four L1s. Studies in Second Language Acquisition, 19(1), 1–16. https://doi.org/10.1017/S0272263197001010 DOI: https://doi.org/10.1017/S0272263197001010
  19. Derwing, T. M., & Munro, M. J. (2009). Putting accent in its place: Rethinking obstacles to communication. Language Teaching, 42(4), 476–490. https://doi.org/10.1017/S026144480800551X DOI: https://doi.org/10.1017/S026144480800551X
  20. Dillon, T., & Wells, D. (2021). Student perceptions of mobile automated speech recognition for pronunciation study and testing. English Teaching, 76(4), 101–122. https://doi.org/10.15858/engtea.76.4.202112.101 DOI: https://doi.org/10.15858/engtea.76.4.202112.101
  21. Dutta, S., Tao, S.A., Reyna, J.C., Hacker, R.E., Irvin, D.W., Buzhardt, J.F., Hansen, J.H.L. (2022). Challenges remain in building ASR for spontaneous preschool children speech in naturalistic educational environments. Proceedings of Interspeech 2022, 4322-4326. https://doi.org/10.21437/Interspeech.2022-555 DOI: https://doi.org/10.21437/Interspeech.2022-555
  22. Eskenazi, M. (1999). Using automatic speech processing for foreign language pronunciation tutoring: Some issues and a prototype. Language Learning & Technology, 2(2), 62–76. https://doi.org/10.64152/10125/25043 DOI: https://doi.org/10.64152/10125/25043
  23. Evanini, K. (2019). Overview of automated scoring. In K. Evanini, & K. Zechner (Eds.), Using language technologies to score spontaneous speech (pp. 3–20). Routledge. DOI: https://doi.org/10.4324/9781315165103-1
  24. Evanini, K., & Wang, X. (2013). Automated speech scoring for non-native middle school students with multiple task types. Interspeech 2013, 2435–2439. https://doi.org/10.21437/Interspeech.2013-566 DOI: https://doi.org/10.21437/Interspeech.2013-566
  25. Evers, K., & Chen, S. (2021). Effects of automatic speech recognition software on pronunciation for adults with different learning styles. Journal of Educational Computing Research, 59(4), 669–685. https://doi.org/10.1177/0735633120972011 DOI: https://doi.org/10.1177/0735633120972011
  26. Feng, S., Kudina, O., Halpern, B. M., & Scharenborg, O. (2021, August 30). Quantifying bias in automatic speech recognition. 22nd Annual Conference of the International Speech Communication Association (INTERSPEECH 2021). International Speech Communication Association, Brno, Czech Republic. http://arxiv.org/abs/2103.15122
  27. Ferland, L., Huffstutler, T., Rice, J., Zheng, J., Ni, S., & Gini, M. (2019). Evaluating older users’ experiences with commercial dialogue systems: Implications for future design and development. arXiv. http://arxiv.org/abs/1902.04393
  28. Filippidou, F., & Moussiades, L. (2020). Α benchmarking of IBM, Google and Wit automatic speech recognition systems. In I. Maglogiannis, L. Iliadis, & E. Pimenidis (Eds.), Artificial Intelligence Applications and Innovations (Vol. 583, pp. 73–82). Springer International Publishing. https://doi.org/10.1007/978-3-030-49161-1_7 DOI: https://doi.org/10.1007/978-3-030-49161-1_7
  29. Fuckner, M., Horsman, S., Wiggers, P., & Janssen, I. (2023). Uncovering bias in ASR systems: Evaluating Wav2vec2 and Whisper for Dutch speakers. 2023 International Conference on Speech Technology and Human-Computer Dialogue (SpeD), 146–151. https://doi.org/10.1109/SpeD59241.2023.10314895 DOI: https://doi.org/10.1109/SpeD59241.2023.10314895
  30. Gao, L., Tejedor-Garcia, C., Strik, H., Cucchiarini, C. (2024) Reading miscue detection in primary school through automatic speech recognition. Proceedings of Interspeech 2024, 5153-5157. https://doi.org/10.21437/Interspeech.2024-1180 DOI: https://doi.org/10.21437/Interspeech.2024-1180
  31. Georgescu, A.-L., Pappalardo, A., Cucu, H., & Blott, M. (2021). Performance vs. Hardware requirements in state-of-the-art automatic speech recognition. EURASIP Journal on Audio, Speech, and Music Processing, 2021(28), 1–30. https://doi.org/10.1186/s13636-021-00217-4 DOI: https://doi.org/10.1186/s13636-021-00217-4
  32. Hartshorne, J. K., Tenenbaum, J. B., & Pinker, S. (2018). A critical period for second language acquisition: Evidence from 2/3 million English speakers. Cognition, 177, 263–277. https://doi.org/10.1016/j.cognition.2018.04.007 DOI: https://doi.org/10.1016/j.cognition.2018.04.007
  33. Hayes, A. F. (2022). Introduction to mediation, moderation, and conditional process analysis: A regression-based approach (3rd ed.). The Guilford Press
  34. Hinsvark, A., Delworth, N., Rio, M. D., McNamara, Q., Dong, J., Westerman, R., Huang, M., Palakapilly, J., Drexler, J., Pirkin, I., Bhandari, N., & Jette, M. (2021). Accented speech recognition: A survey (No. arXiv:2104.10747). arXiv. http://arxiv.org/abs/2104.10747
  35. Hollands, S., Blackburn, D., & Christensen, H. (2022). Evaluating the performance of state-of-the-art ASR systems on non-native English using corpora with extensive language background variation. Interspeech 2022, 3958–3962. https://doi.org/10.21437/Interspeech.2022-10433 DOI: https://doi.org/10.21437/Interspeech.2022-10433
  36. Inceoglu, S., Chen, W.-H., & Lim, H. (2023). Assessment of L2 intelligibility: Comparing L1 listeners and automatic speech recognition. ReCALL, 35(1), 89–104. https://doi.org/10.1017/S0958344022000192 DOI: https://doi.org/10.1017/S0958344022000192
  37. Johnson, C., Cardoso, W., Zuercher, B., Brannen, K., & Springer, S. (2024). Assessing pronunciation using dictation tools: The use of Google Voice Typing to score a pronunciation placement test. Journal of Second Language Pronunciation, 10(1), 10–34. https://doi.org/10.1075/jslp.23033.joh DOI: https://doi.org/10.1075/jslp.23033.joh
  38. Kang, O., & Johnson, D. (2018). The roles of suprasegmental features in predicting English oral proficiency with an automated system. Language Assessment Quarterly, 15(2), 150–168. https://doi.org/10.1080/15434303.2018.1451531 DOI: https://doi.org/10.1080/15434303.2018.1451531
  39. Kang, O., & Rubin, D. L. (2009). Reverse linguistic stereotyping: Measuring the effect of listener expectations on speech evaluation. Journal of Language and Social Psychology, 28(4), 441–456. https://doi.org/10.1177/0261927X09341950 DOI: https://doi.org/10.1177/0261927X09341950
  40. Kathiresan, T. (2021). Gender bias in voice recognition: An i- and x-vector-based gender-specific automatic speaker recognition study. In C. Bernardasci, D. Dipino, D. Garassino, S. Negrinelli, E. Pellegrino, & S. Schmid (Eds.), L’individualità del parlante nelle scienze fonetiche: Applicazioni tecnologiche e forensi (Vol. 8, pp. 113–122). Officinaventuno. https://doi.org/10.17469/O2108AISV000006
  41. Kochem, T., Beck, J., & Goodale, E. (2022). Use of ASR-equipped software in the teaching of suprasegmental features of pronunciation: A critical review. CALICO Journal, 39(3). https://doi.org/10.1558/cj.19033 DOI: https://doi.org/10.1558/cj.19033
  42. Knowles, M. S., Holton, E. F., & Swanson, R. A. (2015). The adult learner: The definitive classic in adult education and human resource development (8th ed.). Routledge.
  43. Koenecke, A., Nam, A., Lake, E., Nudell, J., Quartey, M., Mengesha, Z., Toups, C., Rickford, J. R., Jurafsky, D., & Goel, S. (2020). Racial disparities in automated speech recognition. Proceedings of the National Academy of Sciences, 117(14), 7684–7689. https://doi.org/10.1073/pnas.1915768117 DOI: https://doi.org/10.1073/pnas.1915768117
  44. Krishnan, A., Abdullah, B. M., & Klakow, D. (2024). On the encoding of gender in transformer-based ASR representations. Interspeech 2024, 3090–3094. https://doi.org/10.21437/Interspeech.2024-2209 DOI: https://doi.org/10.21437/Interspeech.2024-2209
  45. Lee, J. Y. (2015). Aging and speech understanding. Journal of Audiology and Otology, 19(1), 7–13. https://doi.org/10.7874/jao.2015.19.1.7 DOI: https://doi.org/10.7874/jao.2015.19.1.7
  46. Levis, J., & Suvorov, R. (2012). Automatic speech recognition (ASR). In C. Chapelle (Ed.), The concise encyclopedia of applied linguistics. John Wiley & Sons. DOI: https://doi.org/10.1002/9781405198431.wbeal0066
  47. Li, J. (2022). Recent advances in end-to-end automatic speech recognition. APSIPA Transactions on Signal and Information Processing, 11(1), 1–64. https://doi.org/10.1561/116.00000050 DOI: https://doi.org/10.1561/116.00000050
  48. Liakin, D., Cardoso, W., & Liakina, N. (2017). Mobilizing instruction in a second-language context: Learners’ perceptions of two speech technologies. Languages, 2(3), 11. https://doi.org/10.3390/languages2030011 DOI: https://doi.org/10.3390/languages2030011
  49. Linville, S. E., & Rens, J. (2001). Vocal tract resonance analysis of aging voice using long-term average spectra. Journal of Voice, 15(3), 323–330. https://doi.org/10.1016/S0892-1997(01)00034-0 DOI: https://doi.org/10.1016/S0892-1997(01)00034-0
  50. Liu, Z., Veliche, I.-E., & Peng, F. (2022). Model-based approach for measuring the fairness in ASR. ICASSP 2022 - 2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 6532–6536. https://doi.org/10.1109/ICASSP43922.2022.9747654 DOI: https://doi.org/10.1109/ICASSP43922.2022.9747654
  51. Madnani, N., Loukina, A., Von Davier, A., Burstein, J., & Cahill, A. (2017). Building better open-source tools to support fairness in automated scoring. Proceedings of the First ACL Workshop on Ethics in Natural Language Processing, 41–52. https://doi.org/10.18653/v1/W17-1605 DOI: https://doi.org/10.18653/v1/W17-1605
  52. McCrocklin, S. (2022). Exploring technologies available for teaching and learning second language pronunciation. Technological Resources for Second Language Pronunciation Learning and Teaching, 3. DOI: https://doi.org/10.5040/9781978729483.ch-1
  53. McCrocklin, S., & Edalatishams, I. (2020). Revisiting popular speech recognition software for ESL speech. TESOL Quarterly, 54(4), 1086–1097. https://doi.org/10.1002/tesq.3006 DOI: https://doi.org/10.1002/tesq.3006
  54. Mehrabi, N., Morstatter, F., Saxena, N., Lerman, K., & Galstyan, A. (2022). A survey on bias and fairness in machine learning. ACM Computing Surveys, 54(6), 1–35. https://doi.org/10.1145/3457607 DOI: https://doi.org/10.1145/3457607
  55. Mroz, A. (2020). Aiming for advanced intelligibility and proficiency using mobile ASR. Journal of Second Language Pronunciation, 6(1), 12–38. https://doi.org/10.1075/jslp.18030.mro DOI: https://doi.org/10.1075/jslp.18030.mro
  56. Munro, M. J., & Derwing, T. M. (1995). Foreign accent, comprehensibility, and intelligibility in the speech of second language learners. Language Learning, 45(1), 73–97. https://doi.org/10.1111/j.1467-1770.1995.tb00963.x DOI: https://doi.org/10.1111/j.1467-1770.1995.tb00963.x
  57. Nelson, C., & Cardoso, W. (2024). Evaluating the effectiveness of Microsoft Transcribe for automating the assessment of pronunciation in language proficiency tests. EuroCALL 2023. CALL for All Languages - Short Papers. https://doi.org/10.4995/EuroCALL2023.2023.17007 DOI: https://doi.org/10.4995/EuroCALL2023.2023.17007
  58. Ngo, T., Hao-Jan Chen, H., & Kuo-Wei Lai, K. (2024). The effectiveness of automatic speech recognition in ESL/EFL pronunciation: A meta-analysis. ReCALL, 36(1), 4–21. https://doi.org/10.1017/S0958344023000113 DOI: https://doi.org/10.1017/S0958344023000113
  59. Ngueajio, M. K., & Washington, G. (2022). Hey ASR system! Why aren’t you more inclusive? Automatic speech recognition systems’ bias and proposed bias mitigation techniques: A literature review. In J. Y. C. Chen, G. Fragomeni, H. Degen, & S. Ntoa (Eds.), HCI International 2022 – Late breaking papers: Interacting with extended reality and artificial intelligence (Vol. 13518, pp. 421–440). Springer Nature Switzerland. https://doi.org/10.1007/978-3-031-21707-4 DOI: https://doi.org/10.1007/978-3-031-21707-4_30
  60. Nickolai, D. (2024). Quantifying the impact of ASR-based instruction: What does the iSpraak platform learner data show? The EuroCALL Review, 31(1), 16–23. https://doi.org/10.4995/eurocall.2024.20221 DOI: https://doi.org/10.4995/eurocall.2024.20221
  61. Pérez Castillejo, S. (2021). Automatic speech recognition: Can you understand me? In T. Beaven, & F. Rosell-Aguilar (Eds.), Innovative language pedagogy report (1st ed., pp. 122–127). Research-publishing.net. https://doi.org/10.14705/rpnet.2021.50.9782490057863 DOI: https://doi.org/10.14705/rpnet.2021.50.1246
  62. Saito, K., Trofimovich, P., Isaacs, T., & Webb, S. (2016). Re-examining phonological and lexical correlates of second language comprehensibility: The role of rater experience. In T. Isaacs, & P. Trofimovich (Eds.), Second language pronunciation assessment: Interdisciplinary perspectives (pp. 141–156). Multilingual Matters. https://doi.org/10.21832/ISAACS6848 DOI: https://doi.org/10.21832/ISAACS6848
  63. Saito, K., Macmillan, K., Kachlicka, M., Kunihara, T., & Minematsu, N. (2023). Automated assessment of second language comprehensibility: Review, training, validation, and generalization studies. Studies in Second Language Acquisition, 45(1), 234–263. https://doi.org/10.1017/S0272263122000080 DOI: https://doi.org/10.1017/S0272263122000080
  64. Sobti, R., Guleria, K., & Kadyan, V. (2024). Comprehensive literature review on children automatic speech recognition system, acoustic linguistic mismatch approaches and challenges. Multimedia Tools and Applications, 83(35), 81933–81995. https://doi.org/10.1007/s11042-024-18753-4 DOI: https://doi.org/10.1007/s11042-024-18753-4
  65. Stolcke, A., & Droppo, J. (2017). Comparing human and machine errors in conversational speech transcription. Interspeech 2017, 137–141. https://doi.org/10.21437/Interspeech.2017-1544 DOI: https://doi.org/10.21437/Interspeech.2017-1544
  66. Tejedor-García, C., Cardeñoso-Payo, V., & Escudero-Mancebo, D. (2021). Automatic speech recognition (ASR) systems applied to pronunciation assessment of L2 Spanish for Japanese speakers. Applied Sciences, 11(15), 6695. https://doi.org/10.3390/app11156695 DOI: https://doi.org/10.3390/app11156695
  67. Urban, E. (2024, October 25). Language support—Speech service—Azure AI services. https://learn.microsoft.com/en-us/azure/ai-services/speech-service/language-support
  68. Vipperla, R., Renals, S., & Frankel, J. (2008). Longitudinal study of ASR performance on ageing voices. Interspeech 2008, 2550–2553. https://doi.org/10.21437/Interspeech.2008-632 DOI: https://doi.org/10.21437/Interspeech.2008-632
  69. Vipperla, R., Renals, S., & Frankel, J. (2010). Ageing Voices: The Effect of Changes in Voice Parameters on ASR Performance. EURASIP Journal on Audio, Speech, and Music Processing, 2010, 1–10. DOI: https://doi.org/10.1155/2010/525783
  70. Werner, L., Huang, G., & Pitts, B. J. (2019). Automated speech recognition systems and older adults: A literature review and synthesis. Proceedings of the Human Factors and Ergonomics Society Annual Meeting, 63(1), 42–46. https://doi.org/10.1177/1071181319631121 DOI: https://doi.org/10.1177/1071181319631121