Skip to content
OPEN ACCESS Jurnal Informatika dan Teknologi Pendidikan

Informatics and Educational Technology Journal

Open access E-ISSN 2777-0680

Trade-off Evaluation: BERT, RoBERTa, and DistilBERT for IMDb Sentiment Analysis on a Tesla T4 Google Colab Free Runtime

Authors

  • Asih Setiyorini Informatika, PJJ Pascasarjana, Universitas Amikom Yogyakarta, 55283, Indonesia
  • Ema Utami Program Doktor Informatika, Universitas Amikom Yogyakarta, 55283, Indonesia

DOI:

https://doi.org/10.59395/vsgxkp03

Keywords:

BERT, DistilBERT, RoBERTa, sentiment analysis, trade-off

Abstract

This study addressed resource-constrained sentiment classification by balancing predictive quality with computational cost and stochastic variability. This study applied a five-stage structured evaluation protocol to compare BERT, RoBERTa, and DistilBERT on the 50,000-review IMDb dataset. The original experiment covered 27 configurations formed by three models, three random seeds, and three maximum sequence lengths. Architecture-level robustness was summarized using mean ± standard deviation across seeds, while one validation-selected configuration per architecture was retained only for confirmatory profiling on a uniform Tesla T4 Google Colab Free runtime. Performance was evaluated using accuracy, macro-F1, AUC-ROC, and MCC, and efficiency was assessed through checkpoint size, inference latency, and peak VRAM. Paired McNemar tests, bootstrap confidence intervals, classical TF-IDF baselines, Pareto analysis, and manual error inspection complemented the comparison. At sequence length 512, mean test macro-F1 was 0.95403 ± 0.00139 for RoBERTa, 0.94016 ± 0.00042 for BERT, and 0.93569 ± 0.00086 for DistilBERT. In the representative confirmatory runs, RoBERTa achieved the highest macro-F1 (0.9516), whereas DistilBERT provided the smallest checkpoint (256.10 MB), lowest BS16 latency (14.86 ms/sample), and lowest training VRAM (2399.80 MB). The results supported scenario-based model selection rather than a single universally optimal model.

Downloads

Download data is not yet available.

References

Areshey, A., & Mathkour, H. (2024). Exploring transformer models for sentiment classification: A comparison of BERT , ROBERTA , ALBERT , DISTILBERT , and XLNET. Expert Systems, 41(11), e13701. https://doi.org/10.1111/exsy.13701

Băroiu, A.-C., & Trăușan-Matu, Ștefan. (2022). Automatic Sarcasm Detection: Systematic Literature Review. Information, 13(8), 399. https://doi.org/10.3390/info13080399

Bashiri, H., & Naderi, H. (2024). Comprehensive review and comparative analysis of transformer models in sentiment analysis. Knowledge and Information Systems, 66(12), 7305–7361. https://doi.org/10.1007/s10115-024-02214-3

Casola, S., Lauriola, I., & Lavelli, A. (2022). Pre-trained transformers: An empirical comparison. Machine Learning with Applications, 9, 100334. https://doi.org/10.1016/j.mlwa.2022.100334

Cassee, N., Agaronian, A., Constantinou, E., Novielli, N., & Serebrenik, A. (2024). Transformers and meta-tokenization in sentiment analysis for software engineering. Empirical Software Engineering, 29(4), 77. https://doi.org/10.1007/s10664-024-10468-2

Chicco, D., & Jurman, G. (2020). The advantages of the Matthews correlation coefficient (MCC) over F1 score and accuracy in binary classification evaluation. BMC Genomics, 21(1), 6. https://doi.org/10.1186/s12864-019-6413-7

Chicco, D., Tötsch, N., & Jurman, G. (2021). The Matthews correlation coefficient (MCC) is more reliable than balanced accuracy, bookmaker informedness, and markedness in two-class confusion matrix evaluation. BioData Mining, 14(1), 13. https://doi.org/10.1186/s13040-021-00244-z

Dehghani, M., Arnab, A., Beyer, L., Vaswani, A., & Tay, Y. (2022). THE EFFICIENCY MISNOMER.

Devlin, J., Chang, M.-W., Lee, K., & Toutanova, K. (2019). BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. Proceedings of the 2019 Conference of the North, 4171–4186. https://doi.org/10.18653/v1/N19-1423

Ganesh, P., Chen, Y., Lou, X., Khan, M. A., Yang, Y., Sajjad, H., Nakov, P., Chen, D., & Winslett, M. (2021). Compressing Large-Scale Transformer-Based Models: A Case Study on BERT. Transactions of the Association for Computational Linguistics, 9, 1061–1080. https://doi.org/10.1162/tacl_a_00413

Herawan, D. F., & Saputri, T. R. D. (2025). Benchmarking Model Transformer Modern untuk Analisis Sentimen dan Tren Konsumen dalam Industri Fashion. Edumatic: Jurnal Pendidikan Informatika, 9(3), 945–954. https://doi.org/10.29408/edumatic.v9i3.32657

Huang, S., Tang, E., Li, S., Ping, X., & Chen, R. (2022). Hardware-friendly compression and hardware acceleration for transformer: A survey. Electronic Research Archive, 30(10), 3755–3785. https://doi.org/10.3934/era.2022192

Islam, M. T., Parvin, F., Sazan, S. A., & Bin Amir, T. (2024). Comparative Analysis of Sentiment Classification on IMDB 50k Movie Reviews: A Study Using CNN, LSTM, CNN-LSTM, and BERT Models. 2024 IEEE International Conference on Power, Electrical, Electronics and Industrial Applications (PEEIACON), 512–517. https://doi.org/10.1109/PEEIACON63629.2024.10800035

Jim, J. R., Talukder, M. A. R., Malakar, P., Kabir, M. M., Nur, K., & Mridha, M. F. (2024). Recent advancements and challenges of NLP-based sentiment analysis: A state-of-the-art review. Natural Language Processing Journal, 6, 100059. https://doi.org/10.1016/j.nlp.2024.100059

Lin, T., Wang, Y., Liu, X., & Qiu, X. (2022). A survey of transformers. AI Open, 3, 111–132. https://doi.org/10.1016/j.aiopen.2022.10.001

Liu, H.-I., Galindo, M., Xie, H., Wong, L.-K., Shuai, H.-H., Li, Y.-H., & Cheng, W.-H. (2024). Lightweight Deep Learning for Resource-Constrained Environments: A Survey. ACM Computing Surveys, 56(10), 1–42. https://doi.org/10.1145/3657282

Mao, Y., Liu, Q., & Zhang, Y. (2024). Sentiment analysis methods, applications, and challenges: A systematic literature review. Journal of King Saud University - Computer and Information Sciences, 36(4), 102048. https://doi.org/10.1016/j.jksuci.2024.102048

Minaee, S., Kalchbrenner, N., Cambria, E., Nikzad, N., Chenaghlu, M., & Gao, J. (2022). Deep Learning--based Text Classification: A Comprehensive Review. ACM Computing Surveys, 54(3), 1–40. https://doi.org/10.1145/3439726

Paneru, B., Thapa, B., & Paneru, B. (2025). Sentiment analysis of movie reviews: A flask application using CNN with RoBERTa embeddings. Systems and Soft Computing, 7, 200192. https://doi.org/10.1016/j.sasc.2025.200192

Papia, S. K., Khan, M. A., Habib, T., Rahman, M., & Islam, M. N. (2024). DistilRoBiLSTMFuse: An efficient hybrid deep learning approach for sentiment analysis. PeerJ Computer Science, 10, e2349. https://doi.org/10.7717/peerj-cs.2349

Puspita, R., & Rahayu, C. (2023). Sentiment Analysis on IMDB Movie Reviews using BERT. Indonesian Journal of Artificial Intelligence and Data Mining, 6(2), 179. https://doi.org/10.24014/ijaidm.v6i2.24239

Sanh, V., Debut, L., Chaumond, J., & Wolf, T. (2019). DistilBERT, a distilled version of BERT: Smaller, faster, cheaper and lighter (Version 4). arXiv. https://doi.org/10.48550/ARXIV.1910.01108

Tay, Y., Dehghani, M., Bahri, D., & Metzler, D. (2023). Efficient Transformers: A Survey. ACM Computing Surveys, 55(6), 1–28. https://doi.org/10.1145/3530811

Wankhade, M., Rao, A. C. S., & Kulkarni, C. (2022). A survey on sentiment analysis methods, applications, and challenges. Artificial Intelligence Review, 55(7), 5731–5780. https://doi.org/10.1007/s10462-022-10144-1

Zhang, H., & Shafiq, M. O. (2024). Survey of transformers and towards ensemble learning using transformers for natural language processing. Journal of Big Data, 11(1), 25. https://doi.org/10.1186/s40537-023-00842-0

Downloads

Published

2026-10-08

How to Cite

Setiyorini, A., & Utami, E. . (2026). Trade-off Evaluation: BERT, RoBERTa, and DistilBERT for IMDb Sentiment Analysis on a Tesla T4 Google Colab Free Runtime. Jurnal Informatika Dan Teknologi Pendidikan, 6(2), 77-86. https://doi.org/10.59395/vsgxkp03

Similar Articles

You may also start an advanced similarity search for this article.