This study presents a controlled empirical benchmark of Python and R for sentiment classification and topic modeling, addressing the ongoing debate regarding the relative suitability of both programming environments for natural language processing tasks. To ensure a fair comparison, identical experimental conditions were established for both platforms. Using the Pang-Lee corpus of 10,662 labeled review snippets, mirrored multinomial Naive Bayes implementations were evaluated by stratified five-fold cross-validation, while matched six-topic Latent Dirichlet Allocation (LDA) models employed identical preprocessing procedures, tokenization, vocabulary construction, hyperparameter settings, initialization strategy, and Gibbs-sampling order. Classification performance was assessed using Accuracy, Precision, Recall, F1-score, and ROC-AUC, whereas topic-model quality was evaluated using UMass coherence and perplexity. Execution performance was measured under identical WebAssembly-based environments to eliminate hardware-related bias. The experimental results demonstrate that Python and R produced identical predictive performance, achieving a mean Accuracy of 0.7459, F1-score of 0.7450, ROC-AUC of 0.8248, UMass coherence of −2.9997, and perplexity of 318.7501. Although model quality remained unchanged across programming languages, substantial differences were observed in execution efficiency, with the median complete-pipeline runtime equal to 0.833 s for Python and 17.220 s for R under the same execution conditions. Based on these findings, a seven-criterion decision framework is proposed that distinguishes predictive performance, topic-model quality, execution efficiency, statistical workflow, visualization capabilities, reproducibility, and deployment considerations. The proposed framework provides practical guidance for researchers and practitioners selecting between Python, R, or hybrid analytical workflows for text mining and natural language processing applications.
P. Joyce, “Python Programming,” in C and Python Applications, pp. 1-57, 2021, [Online]. Available: https://doi.org/10.1007/978-1-4842-7774-4_1.
W. McKinney, Python for Data Analysis: Data Wrangling with Pandas, NumPy, and IPython. O’Reilly Media, 2012, [Online]. Available: https://doi.org/10.5555/2361993.
K. I. Musa, W. N. A. W. Mansor, and T. M. Hanis, “R, RStudio and RStudio Cloud,” in Data Analysis in Medicine and Health Using R, pp. 1-14, 2023, [Online]. Available: https://doi.org/10.1201/9781003296775-1.
C. Sievert, Interactive Web-Based Data Visualization with R, plotly, and shiny. Chapman and Hall/CRC, 2020, [Online]. Available: https://doi.org/10.1201/9780429447273.
M. M. Talipov, “Computational modeling and analysis of mechanical power consumption in train assemblers’ work,” Proceedings of International Conference on Applied Innovation in IT, vol. 13, no. 2, pp. 419-426, 2025, [Online]. Available: https://doi.org/10.25673/120513.
R. Salmorbekova and M. Talipov, “Digital risk matrix and safety management workflow for airport infrastructure in developing countries: a data-driven prioritization approach,” Vibroengineering Procedia, vol. 62, pp. 653-661, Jun. 2026, [Online]. Available: https://doi.org/10.21595/vp.2026.26121.
M. Shukurova, K. Ruziev, M. Talipov, and G. Talipova, “3D modeling of filtration in oil and gas reservoirs based on satellite data,” Proceedings of International Conference on Applied Innovation in IT, vol. 14, no. 1, pp. 267-271, Mar. 2026, [Online]. Available: https://doi.org/10.25673/123569.
M. Shukurova, M. Talipov, K. Ruziev, and K. Jurayeva, “Advanced geospatial monitoring of oil and gas infrastructure via satellite data,” Mathematical Models in Engineering, vol. 12, no. 2, pp. 190-201, Jun. 2026, [Online]. Available: https://doi.org/10.21595/mme.2026.25328.
Y. Croissant and G. Millo, “Panel data econometrics in R: The plm package,” Journal of Statistical Software, vol. 27, no. 2, pp. 1-43, 2008, [Online]. Available: https://doi.org/10.18637/jss.v027.i02.
S. Bird, E. Klein, and E. Loper, Natural Language Processing with Python. O’Reilly Media, 2009, [Online]. Available: https://doi.org/10.5555/1717171.
M. Honnibal and I. Montani, “spaCy 2: Natural language understanding with Bloom embeddings, convolutional neural networks and incremental parsing,” 2017, [Online]. Available: https://doi.org/10.5281/zenodo.1212303.
K. Benoit et al., “quanteda: An R package for the quantitative analysis of textual data,” Journal of Open Source Software, vol. 3, no. 30, p. 774, 2018, [Online]. Available: https://doi.org/10.21105/joss.00774.
I. Feinerer, K. Hornik, and D. Meyer, “Text mining infrastructure in R,” Journal of Statistical Software, vol. 25, no. 5, pp. 1-54, 2008, [Online]. Available: https://doi.org/10.18637/jss.v025.i05.
Y. Xie, J. J. Allaire, and G. Grolemund, R Markdown: The Definitive Guide. Boca Raton, FL, USA: Chapman and Hall/CRC, 2018, [Online]. Available: https://bookdown.org/yihui/rmarkdown/.
T. Kluyver et al., “Jupyter Notebooks – a publishing format for reproducible computational workflows,” in F. Loizides and B. Schmidt, Eds., Positioning and Power in Academic Publishing: Players, Agents and Agendas. IOS Press, 2016, pp. 87-90, [Online]. Available: https://doi.org/10.3233/978-1-61499-649-1-87.
P. Joyce, “Embedded Python,” in C and Python Applications, pp. 151-181, 2021, [Online]. Available: https://doi.org/10.1007/978-1-4842-7774-4_5.
D. M. Blei, A. Y. Ng, and M. I. Jordan, “Latent Dirichlet allocation,” Journal of Machine Learning Research, vol. 3, pp. 993-1022, 2003, [Online]. Available: https://doi.org/10.1162/jmlr.2003.3.4-5.993.
B. Pang and L. Lee, “Seeing Stars: Exploiting Class Relationships for Sentiment Categorization with Respect to Rating Scales,” in Proceedings of the 43rd Annual Meeting of the Association for Computational Linguistics (ACL’05), Ann Arbor, MI, pp. 115-124, 2005, [Online]. Available: https://doi.org/10.3115/1219840.1219855.
Pyodide Developers, “Pyodide: Python distribution for the browser and Node.js based on WebAssembly,” Version 314.0.2, 2026, [Online]. Available: https://pyodide.org/, [Accessed: Jul. 8, 2026].
R-wasm Project, “webR - R in the Browser,” Version 0.6.0, 2026, [Online]. Available: https://docs.r-wasm.org/, [Accessed: Jul. 8, 2026].