The modality of speech sample presentation constitutes a methodological challenge for researchers who seek to understand the impact of non-phonological variables on (second language) L2 speech comprehensibility. This pilot-study aims to address this challenge. It uses a mixed-method approach to explore the possibility offered by text-to-speech software to present speech samples in an oral modality while also controlling for phonological variables. We generated audio files based on verbatim transcription (i.e., including run-on sentences, filler words, self-repetitions, colloquialisms) of two extemporaneous speech samples produced in Quebec French, with the help of a commercial text-to-speech software. In addition to the original transcripts, we created two versions of each sample containing decreasingly fewer features typical of spontaneous speech, since it is unclear to what extent such features distract listeners’ attention away from non-phonological features in French. Seven participants then rated the comprehensibility of the audio files on a scale from 1 to 9. They also shared their thought process through the think-aloud method and responded to a retrospective interview question formulated to elicit listener attitudes to artificial speech. The main findings were that speech samples read by an artificial voice were perceived as relatively easy to understand when the transcripts had been cleaned from filled pauses, repetitions and reformulations, and that listeners generally held relatively positive attitudes to AI-generated speech. These findings suggest that text-to-speech could be a useful method for presenting speech samples in L2 comprehensibility research, ultimately contributing to a deeper understanding of the interplay between linguistic factors and comprehensibility in second language acquisition.