Comparing Linguistic Features of AI-Generated and Students-Produced Essays: A Corpus-Based Comparison
Abstract
The integration of artificial intelligence (AI) into language teaching and learning has concerned language teachers especially in distinguishing the content produced by AI and by learners. Previous studies commonly emphasised on the grammatical and lexical aspects; however, less attention is given to research specifically examining on discourse markers which are essential for textual cohesion and coherence. Thus, this corpus-based study is conducted to compare the use of discourse markers in academic essays produced by AI and student writers. Two comparable corpora are created: the first corpus consists of 44 essays generated by AI and another corpus comprises of 44 essays written by university students without any support from AI. The LancsBox X was used to analyse the frequency, variety, and functional distribution of discourse markers categories (additive/elaborative, contrastive, causal, and sequential). Results indicate a significant difference in discourse markers use between the two groups. AI-generated texts are found to exhibit over-reliance on elaborative markers, meanwhile, inferential and temporal markers are found more frequently in the student essays. These findings suggest that although the academic essays produced by AI are structurally flawless, it lacks the nuanced, stylistic and argumentative depth found in the students’ essay.
Downloads
References
Abdullah, S. (2015). A data-driven contrastive study on Malay ESL learners' use of lexical verbs and verb-noun collocations in argumentative writing (Doctoral dissertation, Universiti Teknologi MARA, Shah Alam, Malaysia). UiTM Institutional Repository. https://ir.uitm.edu.my/id/eprint/19589/
Adnan, A. H. M. (2023). Corpus-based analysis of lexical errors in paragraph writing among low-proficiency Malaysian ESL diploma students. Journal of Creative Practices in Language Learning and Teaching, 11(2), 45–61. https://doi.org/10.24191/jcpilt.v11i2.23412
Amirjalili, F., Neysani, M., & Nikbakht, A. (2024). Exploring the boundaries of authorship: A comparative analysis of AI-generated text and human academic writing in English literature. Frontiers in Education, 9, Article 1347421. https://doi.org/10.3389/feduc.2024.1347421
Biber, D. (1993). Representativeness in corpus design. Literary and Linguistic Computing, 8(4), 243–257. https://doi.org/10.1093/llc/8.4.243
Biber, D., Conrad, S., & Reppen, R. (1998). Corpus linguistics: Investigating language structure and use. Cambridge University Press.
Brezina, V., & Platt, W. (2025). #LancsBox X [Computer software]. Lancaster University. http://lancsbox.lancs.ac.uk
Conde, J., Reviriego, P., Salvachúa, J., Martínez, G., Hernández, J. A., & Lombardi, F. (2024). Understanding the impact of artificial intelligence in academic writing: Metadata to the rescue. Computer, 57(1), 105–109. https://doi.org/10.1109/mc.2023.3327330
Crossley, S. A., Salsbury, T., & McNamara, D. S. (2014). Assessing lexical proficiency using analytic ratings: A case for collocation accuracy. Applied Linguistics, 36(5), 570–590. https://doi.org/10.1093/applin/amt056
Erliani, Nys. N. P., Eryansyah, & Amrullah. (2025). The Impact of AI Tools on EFL Students’ Self-Efficacy in Academic Writing: A Qualitative Study. Edukasi: Jurnal Pendidikan Dan Pengajaran, 12(01), 6–17. https://doi.org/10.19109/s75mbm13
Fredrick, D. R., & Craven, L. (2025). Lexical diversity, syntactic complexity, and readability: A corpus-based analysis of ChatGPT and L2 student essays. Frontiers in Education, 10, Article 1616935. https://doi.org/10.3389/feduc.2025.1616935
Georgiou, G. P. (2024). Differentiating between human-written and AI-generated texts using linguistic features automatically extracted from an online computational tool. arXiv. https://doi.org/10.48550/arxiv.2407.03646
Goyibova, N., Muslimov, N., Kannazarova, Z., Kadirova, N., Alautdinova, K., & Ismatullaeva, I. (2025). Exploring the impact of artificial intelligence on academic writing: A bibliometric analysis of trends, advancements, and ethical challenges. Forum for Linguistic Studies, 7(6), 342–360. https://doi.org/10.30564/fls.v7i6.9054
Herbold, S., Hautli-Janisz, A., Heuer, U., Kikteva, Z., & Trautsch, A. (2023). A large-scale comparison of human-written versus ChatGPT-generated essays. Scientific Reports, 13(1), Article 18617. https://doi.org/10.1038/s41598-023-45644-9
Herbold, S., Hautli-Janisz, A., Heuer, U., Krell, M. T., Lorbach, J., Dey, M., Zhang, B., Schwinn, A., Trautsch, L., Lütjen, B., & Heidinger, C. (2023). A large-scale comparison of human-written and AI-generated essays. Scientific Data, 10(1), 1–15. https://doi.org/10.1038/s41597-023-02301-w
Hyland, K. (2005). Metadiscourse: Exploring interaction in writing. Continuum.
Jen, S. L., & Salam, A. R. (2024). Using artificial intelligence for essay writing. Arab World English Journal, (Special Issue on ChatGPT), 90–99. https://doi.org/10.24093/awej/ChatGPT.5
Jiang, Y. (2024). Interaction and dialogue: Integration and application of artificial intelligence in blended mode writing feedback. The Internet and Higher Education, 64, 100975–100975. https://doi.org/10.1016/j.iheduc.2024.100975
Jiang, F. K., & Hyland, K. (2024). Framing student writing: Metadiscourse in human and AI-generated essays. Assessing Writing, 62, Article 100863. https://doi.org/10.1016/j.asw.2024.100863
Kamaludin, P. N. H. B. (2014). Factors affecting the assessment of ESL students' writing: A case study of UiTM English language lecturers (Unpublished master's dissertation). Universiti Teknologi MARA.
Lin, Z. (2023, October 19). Techniques for supercharging academic writing with generative AI. PsyArXiv. https://doi.org/10.31234/osf.io/9yhwz
Liu, Y., Han, T., Sun, S., Liang, C., & Passonneau, R. J. (2023). ArguGPT: Exploiting LLMs for generating argumentative essays and analyzing their stylistic signatures. arXiv. https://doi.org/10.48550/arXiv.2304.14072
Lu, X. (2010). Automatic analysis of syntactic complexity in second language writing. International Journal of Corpus Linguistics, 15(4), 474–496. https://doi.org/10.1075/ijcl.15.4.02lu
Mat Zali, M., Mohd Razlan, R., Raja Baniamin, R. M., & Setia, R. (2022). Interactional metadiscourse analysis of ESL learners’ essays. Environment-Behaviour Proceedings Journal, 7(SI 9), 55–60. https://doi.org/10.21834/ebpj.v7isi9.4248
Mizumoto, A., & Eguchi, M. (2023). Exploring the linguistic fingerprints of AI-generated academic texts: A corpus-driven comparison with human scholarship. Journal of English for Academic Purposes, 65, Article 101287. https://doi.org/10.1016/j.jeap.2023.101287
Mohamed, A. F., Ab Rashid, R., Lateh, N. H. M., & Kurniawan, Y. (2021). The use of metadiscourse in good Malaysian undergraduate persuasive essays. INSANIAH: Online Journal of Language, Communication, and Humanities, 4(1), 1–12.
Ramachandran, S., Jalaluddin, I., & Sharatol Ahmad Shah, S. (2024). Investigating lexical collocational errors in the argumentative essays of Malaysian tertiary ESL learners. 3L: Language, Linguistics, Literature, 30(1), 112–128. https://doi.org/10.17576/3L-2024-3001-08
Shi, J., Liu, W., & Hu, K. (2025). Exploring how AI literacy and self-regulated learning relate to student writing performance and well-being in generative AI-supported higher education. Behavioral Sciences, 15(5), Article 705. https://doi.org/10.3390/bs15050705
Ting, S. H., Raslie, H., & Jee, L. J. (2011). Case study on the persuasiveness of argument texts written by proficient and less proficient Malaysian undergraduates. Malaysian Journal of Learning and Instruction, 8, 71–92. https://doi.org/10.32890/mjli.8.2011.7627
Yildiz Durak, H., Eğin, F., & Onan, A. (2025). A comparison of human-written versus AI-generated text in discussions at educational settings: Investigating features for ChatGPT, Gemini and BingAI. European Journal of Education, 60, Article e70014. https://doi.org/10.1111/ejed.70014
Zaheer, S., Ma, C., Zhu, Y., & Vasinda, S. (2025). GenAI in academic writing—empowering learners or redefining traditional pedagogical practices?: A systematic review from 2019-2023. International Journal of Artificial Intelligence (AI) in Teaching and Learning (IJAITL), 1(1), 1–34. https://doi.org/10.4018/IJAITL.373582















