A Small-Scale Comparative Evaluation of ChatGPT-3.5 and Google Translate for English-Urdu and Urdu-English Translation

Authors

  • Aatka Faryal Riaz Department of Computer Science, International University of Applied Science, Bad Honnef, Germany
  • Rafay Azmat Department of Electronics Engineering, Mehran University of Engineering & Technology, Jamshoro, Pakistan
  • Rubeka Sahar ICT Department, University of Sindh, Jamshoro, Pakistan

DOI:

https://doi.org/10.53560/PPASA(63-1)711

Keywords:

Machine Translators, Chat GPT, Google Translate, English-Urdu Language, Urdu-English Language, Evaluation Metrics, BLEU, METEOR

Abstract

Today, technologies like ChatGPT and Google Translate are utilized for translating text and enabling smooth exchange of information. Urdu is the national language of Pakistan, while English serves as the official language and the primary medium of communication in schools and workplaces. Since people commonly speak Urdu, they often use machine translators for English to Urdu and Urdu to English translation. These machine translators can produce translations with a diverse vocabulary but sometimes their translations are not contextually correct or fully reliable. Therefore, it is essential to evaluate the translation outputs of the machine translators. The prime objective of this paper is to compare the performance of machine translation systems at the paragraph level, enabling users to identify the most effective one. In this research, Google Translate and ChatGPT-3.5 were used as translation tools to translate text between English and Urdu in both directions. The texts were selected from different categories, such as historical texts, poetry, literary works, novels, and legal documents.  The machine translation systems translated English text into Urdu, and the generated outputs were compared and assessed by authors. Similarly, Urdu text was translated into English, and the resulting translations were compared and assessed by authors. We assessed and compared the translated outputs based on phrasing style, linguistic accuracy, and contextual meaning. To ensure a reliable comparison, we calculated the BiLingual Evaluation Understudy (BLEU) and Metric for Evaluation of Translation with Explicit ORdering (METEOR) scores of the translations generated by these translation tools. Based on our analysis and BLEU and METEOR scores, we concluded that the performance of ChatGPT-3.5 is better than Google Translate.

References

1. Ethnologue. The Ethnologue 200 (2025). https://www.ethnologue.com/insights/ethnologue200/

2. International Center for Language Studies. Most Spoken Languages in the World (2024). https://www.icls.edu/blog/most-spoken-languages-in-the-world/

3. A.A. Malik and A. Habib. Urdu to English machine translation using bilingual evaluation understudy. International Journal of Computer Applications 82(7): 5-12 (2013). https://research.ijcaonline.org/volume82/number7/pxc3891040.pdf

4. M. Ghassemiazghandi. An evaluation of ChatGPT's translation accuracy using BLEU score. Theory and Practice in Language Studies 14(4): 985-994 (2024). https://tpls.academypublication.com/index.php/tpls/article/view/7867/6371

5. S.S. Biswas. Potential use of ChatGPT in global warming. Annals of Biomedical Engineering 51(6): 1126-1127 (2023). https://doi.org/10.1007/s10439-023-03171-8

6. S.S. Biswas. Role of ChatGPT in public health. Annals of Biomedical Engineering 51(5): 868-869 (2023). https://link.springer.com/article/10.1007/s10439-023-03172-7

7. P.P. Ray. ChatGPT: A comprehensive review on background, applications, key challenges, bias, ethics, limitations and future scope. Internet of Things and Cyber-Physical Systems 3: 121-154 (2023).

https://doi.org/10.1016/j.iotcps.2023.04.003

8. A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A.N. Gomez, L. Kaiser, and I. Polosukhin. Attention is all you need. Advances in Neural Information Processing Systems 30: (2017). https://doi.org/10.48550/arXiv.1706.03762

9. P. Memon. Discover how ChatGPT is trained! (2023). https://www.linkedin.com/pulse/discover-how-chatgpt-istrained-pradeep-menon

10. S.C. Siu. ChatGPT and GPT-4 for professional translators: Exploring the potential of large language models in translation. SSRN (2023). https://dx.doi.org/10.2139/ssrn.4448091

11. C. Boitet, H. Blanchon, M. Seligman, and V. Bellynck. Evolution of MT with the Web. Proceedings of the International Conference “Machine Translation 25 Years On”, 21-22 November 2009, Cranfield, England (2009).

https://lig-membres.imag.fr/blanchon/Pdfs/MT25YO-Evolution-MT-web.pdf?utm_source=chatgpt.com

12. I. Caswell. Google Translate learns 24 new languages. Online google blog (2022). https://blog.google/products/translate/24-new-languages/

13. Google Translate. Google Translate Wikipedia (2025). https://en.wikipedia.org/wiki/Google_Translate

14. P. Chang, P.J. Chen, and L.L. Lai. Recursive editing with Google Translate: the impact on writing and error correction. Computer Assisted Language Learning 37(7): 2116-2141 (2024).

https://doi.org/10.1080/09588221.2022.2147192

15. S.C. Tsai. Using Google Translate in EFL drafts: A preliminary investigation. Computer Assisted Language Learning 32(5-6): 510-526 (2019). https://doi.org/10.1080/09588221.2018.1527361

16. K.M. King. Can Google Translate be taught to translate literature? A case for humanists to collaborate in the future of machine translation. Translation Review 105(1): 76-92 (2019).

https://doi.org/10.1080/07374836.2019.1673268

17. Y. Wu, M. Schuster, Z. Chen, Q.V. Le, M.N. Yonghui, W. Macherey, M. Krikun, Y. Cao, Q. Gao, K. Macherey, J. Klingner, A. Shah, M. Johnson, X. Liu, L. Kaiser, S. Gouws, Y. Kato, T. Kudo, H. Kazawa, K. Stevens, G. Kurian, N. Patil, W. Wang, C. Young, J. Smith, J. Riesa, A. Rudnick, O. Vinyals, G. Corrado, M. Hughes, and J. Dean. Google’s neural machine translation system: Bridging the gap between human and machine translation. arXiv preprint arXiv: 1609.08144 (2016). https://doi.org/10.48550/arXiv.1609.08144

18. K. Papineni, S. Roukos, T. Ward, and W.J. Zhu. BLEU: A method for automatic evaluation of machine translation. Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics, Philadelphia, USA (2002). https://aclanthology.org/P02-1040.pdf

19. M.A. Ayesha, S. Noor, M. Ramzan, H.U. Khan, and M. Shoaib. Evaluating Urdu to Arabic machine translation tools. International Journal of Advanced Computer Science and Applications 8(10): 90-96 (2017). https://thesai.org/Downloads/Volume8No10/Paper_12-Evaluating_Urdu_to_Arabic_Machine_Translation_Tools.pdf

20. S. Banerjee and A. Lavie. METEOR: An automatic metric for MT evaluation with improved correlation with human judgments. Proceedings of the ACL Workshop on Intrinsic and Extrinsic Evaluation Measures for Machine Translation and/or Summarization, Ann Arbor, Michigan, USA (2005).

https://aclanthology.org/W05-0909.pdf?utm_source=chatgpt.com

21. S. Khan, N. Aftab, A. Ali, and R. Mohammad. Mutalia Pakistan, Class 12. Punjab Textbook Board, Lahore, Pakistan (2005).

https://www.taleem360.com/12th-class-pak-study-new-punjab-text-book-2022-zqx

22. H. Lamb (Ed.). Genghis Khan, the Emperor of All Men. Doubleday, New York, USA (1927).

https://archive.org/details/genghiskhanemper0000lamb_f6b1/page/n13/mode/2up

Downloads

Published

2026-03-25

How to Cite

Aatka Faryal Riaz, Rafay Azmat, & Rubeka Sahar. (2026). A Small-Scale Comparative Evaluation of ChatGPT-3.5 and Google Translate for English-Urdu and Urdu-English Translation. Proceedings of the Pakistan Academy of Sciences: A. Physical and Computational Sciences, 63(1), 71–81. https://doi.org/10.53560/PPASA(63-1)711

Issue

Section

Research Articles

Similar Articles

1 2 3 4 5 6 7 8 9 > >> 

You may also start an advanced similarity search for this article.