Article Open Access

NLP-Based Prediction of EFL Writing Proficiency: An Interdisciplinary Study Using the ELLIPSE Dataset

(1) * Muhammad Akil Hi Umar Mail (Universitas Persatuan Islam, Indonesia)
*Corresponding author

Abstract


ABSTRACT

This study presents an interdisciplinary collaboration between the English Language Education and Informatics Study Programs at Universitas Persatuan Islam (UNIPI) Bandung, aimed at developing a Natural Language Processing (NLP)-based system for predicting writing proficiency of English as a Foreign Language (EFL) learners. The ELLIPSE corpus—comprising approximately 7,000 essays by 8th-to-12th-grade students, analytically scored across six linguistic dimensions: cohesion, syntax, vocabulary, phraseology, grammar, and conventions—served as the primary dataset. A multi-model approach was employed, comparing classical methods (Ridge Regression, Random Forest, XGBoost) and transformer-based models (DeBERTa-base and DeBERTa-large). The fine-tuned DeBERTa-large model achieved the best performance with a Mean Column-wise RMSE (MCRMSE) of 0.408 and an averaperge R² of 0.768, outperforming all classical baselines. Results reveal that conventions and syntax exhibit the highest predictability, while phraseology and vocabulary demonstrate the greatest scoring variability. Feature importance analysis confirms that type-token ratio, syntactic complexity, and discourse connective usage are the strongest predictors of writing proficiency. Fairness analysis across demographic subgroups—grade level, gender, and economic status—reveals modest but non-trivial prediction disparities warranting attention prior to deployment. Findings contribute to the development of Automated Writing Evaluation (AWE) systems applicable to English language instruction in Indonesian higher education contexts.

Keywords: Automated Writing Evaluation; DeBERTa; ELLIPSE Corpus; English as a Foreign Language; Machine Learning; Natural Language Processing; Writing Proficiency

   

DOI

https://doi.org/10.33122/ejeset.v%25vi%25i.1423
      

Article metrics

Abstract views : 0

   

Cite

   

References


Baker, R. S., & Hawn, A. (2022). Algorithmic bias in education. International Journal of Artificial Intelligence in Education, 32(4), 1052–1092. https://doi.org/10.1007/s40593-021-00285-9

Barrot, J. S., & Agdeppa, J. Y. (2021). Complexity, accuracy, and fluency as indices of college-level L2 writers’ proficiency. Assessing Writing, 47, 100511. https://doi.org/10.1016/j.asw.2020.100511

Crossley, S. A., Tian, Y., Wan, H., & McNamara, D. S. (2023). The ELLIPSE corpus: English Language Learner Insight, Proficiency and Skills Evaluation. In Proceedings of the 18th Workshop on Innovative Use of NLP for Building Educational Applications (BEA) (pp. 1–11). Association for Computational Linguistics. https://doi.org/10.18653/v1/2023.bea-1.1

Devlin, J., Chang, M.-W., Lee, K., & Toutanova, K. (2019). BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of NAACL-HLT 2019 (pp. 4171–4186). Association for Computational Linguistics. https://doi.org/10.18653/v1/N19-1423

Dubey, P., Dubey, P., Raja, R., & Kshatri, S. S. (2025). Bridging language gaps: The role of NLP and speech recognition in oral English instruction. MethodsX, 14, 103359. https://doi.org/10.1016/j.mex.2025.103359

Duarte, J. M., & Berton, L. (2023). A review of semi-supervised learning for text classification. Artificial Intelligence Review, 56, 9401–9469. https://doi.org/10.1007/s10462-023-10393-8

Hartono, W. J., Nurfitri, Ridwan, Kase, E. B. S., Lake, F., & Zebua, R. S. Y. (2023). Artificial Intelligence (AI) solutions in English Language Teaching: Teachers-students perceptions and experiences. Journal on Education, 6(1), 1452–1461. https://jonedu.org/index.php/joe/article/view/3101

He, P., Gao, J., & Chen, W. (2021). DeBERTaV3: Improving DeBERTa using ELECTRA-style pre-training with gradient-disentangled embedding sharing. arXiv preprint arXiv:2111.09543. https://doi.org/10.48550/arXiv.2111.09543

He, P., Liu, X., Gao, J., & Chen, W. (2021). DeBERTa: Decoding-enhanced BERT with disentangled attention. In Proceedings of ICLR 2021. OpenReview. https://doi.org/10.48550/arXiv.2006.03654

Idowu, A. A. (2024). Investigating algorithmic bias in student progress monitoring. Computers and Education: Artificial Intelligence, 7, 100247. https://doi.org/10.1016/j.caeai.2024.100247

Irwanto, I. (2023). Attitudes toward e-learning of undergraduate students during COVID-19: Dataset from Indonesia. Data in Brief, 49, 109380. https://doi.org/10.1016/j.dib.2023.109380

Ke, Z., & Ng, V. (2019). Automated Essay Scoring: A survey of the state of the art. In Proceedings of IJCAI 2019 (pp. 6300–6308). IJCAI. https://doi.org/10.24963/ijcai.2019/879

Knill, K., et al. (2025). Speak & Improve Corpus 2025: An L2 English speech corpus for language assessment and feedback. arXiv preprint arXiv:2412.11986. https://doi.org/10.48550/arXiv.2412.11986

Kumar, M., Singh, N., Wadhwa, J., Singh, P., Kumar, G., & Qtaishat, A. (2024). Utilizing Random Forest and XGBoost data mining algorithms for anticipating students’ academic performance. International Journal of Modern Education and Computer Science, 16(2), 29–44. https://doi.org/10.5815/ijmecs.2024.02.03

Lu, C. (2025). AI-generated corpus learning and EFL learners’ learning of grammatical structures, lexical bundles, and willingness to write. PLOS ONE. https://doi.org/10.1371/journal.pone.0321544

Lyu, J., Chishti, M. I., & Peng, Z. (2022). Marked distinctions in syntactic complexity: A case of second language university learners’ and native speakers’ syntactic constructions. Frontiers in Psychology, 13, 1048286. https://doi.org/10.3389/fpsyg.2022.1048286

Mulyono, H., & Saskia, R. (2020). Dataset on the effects of self-confidence, motivation and anxiety on Indonesian students’ willingness to communicate in face-to-face and digital settings. Data in Brief, 31, 105774. https://doi.org/10.1016/j.dib.2020.105774

Ramesh, D., & Sanampudi, S. K. (2022). An automated essay scoring systems: A systematic literature review. Artificial Intelligence Review, 55(3), 2495–2527. https://doi.org/10.1007/s10462-021-10068-2

Shi, H., & Aryadoust, V. (2023). A systematic review of automated writing evaluation systems. Education and Information Technologies, 28(1), 771–795. https://doi.org/10.1007/s10639-022-11110-8

Slamet, T. S., & Umar, M. A. H. (2025). Critical review of the Technology Acceptance Model in information systems research. International Journal of Strategic Information Systems and Knowledge Management.

Sudarsono, M. A., Kristiani, K., & Setyowibowo, F. (2023). Pengaruh Application Self Efficacy dan Kompleksitas Teknologi terhadap Adopsi Learning Management System oleh Dosen. Journal of Education, 6(1), 2339–2351.

Umar, M. A. H., & Sitohang, B. (2024). Analisis Faktor-Faktor yang Memengaruhi Keputusan Pembelian Paket Wisata Menggunakan Model Klasifikasi Decision Trees, Random Forest dan K-Nearest Neighbours. Journal of Social and Economics Research, 6(2), 25–39.

Umar, M. A. H., & Sitohang, B. (2025). Integrating experience and complexity into the Technology Acceptance Model for health information systems. bit-Tech, 8(2), 1732–1740.

Uto, Y. (2021). A review of deep-neural automated essay scoring models. Behaviormetrika, 48(2), 459–484. https://doi.org/10.1007/s41237-021-00142-y

Vajjala, T., & Lõo, I. (2014). On the applicability of second language acquisition theory for automated essay scoring of EFL student writing. In Proceedings of the Workshop on NLP for Learning (pp. 1–9). Association for Computational Linguistics.

Wang, Y., Wang, C., Li, R., & Lin, H. (2022). On the use of BERT for automated essay scoring: Joint learning of multi-scale essay representation. In Proceedings of NAACL-HLT 2022 (pp. 3416–3425). Association for Computational Linguistics. https://doi.org/10.18653/v1/2022.naacl-main.249

Zhao, R., Zhuang, Y., Zou, D., Quan, X., & Yu, P. L. H. (2022). AI-assisted automated scoring of picture-cued writing tasks for language assessment. Education and Information Technologies, 28(6), 7031–7063. https://doi.org/10.1007/s10639-022-11473-y


Refbacks

  • There are currently no refbacks.


Creative Commons License
This work is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License.

 
This work is licensed under a Creative Commons Attribution-ShareAlike 4.0