Deep Learning Techniques for Image Recognition (Machine Learning)

Kolapo Obanewa, Olumide Innocent Olope

Abstract

Deep learning (DL), a sophisticated subset of machine learning (ML), has emerged as a transformative force within the broader realm of artificial intelligence (AI). By leveraging architectures such as convolutional neural networks (CNNs), DL has significantly advanced image recognition capabilities, enabling systems to identify and classify visual data with remarkable precision accurately. This technology is not only applicable to image recognition. Still, it has also made strides in diverse areas, such as speech recognition, language translation, automated gameplay, healthcare diagnostics, and the development of self-driving vehicles. The success of DL in this domain can be attributed to its ability to learn hierarchical representations of data, allowing for improved feature extraction and pattern recognition. Despite its impressive performance, deep learning is not without its limitations. Key challenges include its reliance on vast amounts of labelled data, which can be difficult and expensive to obtain, its lack of common sense reasoning and difficulties in addressing complex, multifaceted problems.

Additionally, DL models often struggle with long-term planning and decision-making, which can hinder their effectiveness in certain applications. This paper delves into the significant role of deep learning in image recognition, providing a comprehensive overview of its methodologies, applications, strengths, and limitations. By examining current advancements and ongoing challenges, this work aims to contribute to understanding deep learning's impact on the field and its future potential.




Keywords


Deep Learning; Machine Learning; Image Recognition; Data Require-ments; Interpretability; Computational Resources; Overfitting; Adver-sarial Attacks

Full Text:

PDF


References


1. Lecun, Y., Bottou, L., Bengio, Y., & Haffner, P. (1998). Gradient-based learning applied to document recognition. Proceedings of the IEEE, 86(11), 2278–2324. doi: 10.1109/5.726791

2. Ivakhnenko, A. G. (1971). Polynomial Theory of Complex Systems. IEEE Transactions on Systems, Man, and Cybernetics, SMC-5(4), 364–378.

3. Krizhevsky, A., Sutskever, I., & Hinton, G. E. (2012). ImageNet classification with deep convolutional neural networks. Retrieved from https://proceedings.neurips.cc/paper_files/paper/2012/file/c399862d3b9d6b76c8436e924a68c45b-Paper.pdf

4. Esteva, A., Kuprel, B., Novoa, R. A., Ko, J., Swetter, S. M., Blau, H. M., & Thrun, S. (2017). Dermatologist-level classification of skin cancer with deep neural networks. Nature, 542(7639), 115–118. doi: 10.1038/nature21056

5. Marcus, G. (2018). Deep Learning: A Critical Appraisal. Retrieved from https://arxiv.org/abs/1801.00631

6. Hinton, G. E., Osindero, S., & Teh, Y. W. (2006). A fast learning algorithm for deep belief nets. Neural Computation, 18(7), 1527–1554.

7. Simonyan, K., & Zisserman, A. (2014). Very deep convolutional networks for large-scale image recognition. Retrieved from https://arxiv.org/abs/1409.1556

8. He, K., Zhang, X., Ren, S., & Sun, J. (2016). Deep residual learning for image recognition. Retrieved from https://arxiv.org/abs/1512.03385

9. Gulshan, V., Peng, L., Coram, M., Stumpe, M. C., Wu, D., Narayanaswamy, A., Venugopalan, S., Widner, K., Madams, T., Cuadros, J., Kim, R., Raman, R., Nelson, P. C., Mega, J. L., & Webster, D. R. (2016). Development and Validation of a Deep Learning Algorithm for Detection of Diabetic Retinopathy in Retinal Fundus Photographs. JAMA, 316(22), 2402. doi: 10.1001/jama.2016.17216

10. Lin, H.-Y., Chang, C.-K., & Tran, V. L. (2024). Lane detection networks based on deep neural networks and temporal information. Alexandria Engineering Journal, 98, 10–18. doi: 10.1016/j.aej.2024.04.027

11. Wang, L., Xiong, Y., Wang, Z., Qiao, Y., Lin, D., Tang, X., & Gool, L. V. (2016). Temporal segment networks for action recognition in videos. Retrieved from https://arxiv.org/abs/1705.02953

12. Krawczyk, B. (2016). Learning from imbalanced data: open challenges and future directions. Progress in Artificial Intelligence, 5(4), 221–232. doi: 10.1007/s13748-016-0094-0

13. Garsia, V., & Bruna, J. (2017). Few-shot learning with graph neural networks. Retrieved from https://arxiv.org/abs/1711.04043

14. Doshi-Velez, F., & Kim, P. (2017). Towards a rigorous science of interpretable machine learning. Retrieved from https://arxiv.org/abs/1702.08608

15. Gilpin, L., Bau, D., Yuan, B., Bajwa, A., Specter, M., & Kagal, L. (2018). Explaining explanations: An overview of interpretability of machine learning. Retrieved from https://arxiv.org/abs/1806.00069

16. d’Avila Garcez, A. S., Broda, K. B., & Gabbay, D. M. (2002). Neural-Symbolic Learning Systems. In Perspectives in Neural Computing. Springer London. doi: 10.1007/978-1-4471-0211-3

17. Zhou, T., Brown, M., Snavely, N., & Lowe, D. (2020). Unsupervised learning of depth and ego-motion from video. Retrieved from https://arxiv.org/abs/1704.07813

18. Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., Petersen, S., Beattie, C., Sadik, A., Antonoglou, I., King, H., Kumaran, D., Wierstra, D., Legg, S., & Hassabis, D. (2015). Human-level control through deep reinforcement learning. Nature, 518(7540), 529–533. doi: 10.1038/nature14236

19. Amodei, D., Olah, C., Steinhardt, J., Christiano, P., Schulman, J., & Mané, D. (2016). Concrete Problems in AI Safety. Retrieved from https://arxiv.org/abs/1606.06565

20. Zhang, C., Bengio, S., Hardt, M., Recht, B., & Vinyals, O. (2016). Understanding Deep Learning Requires Rethinking Generalization. Retrieved from https://arxiv.org/abs/1611.03530

21. Goodfellow, I., Shlens, J., & Szegedy, C. (2014). Explaining and Harnessing Adversarial Examples. Retrieved from https://arxiv.org/abs/1412.6572

22. Pan, S. J., & Yang, Q. (2010). A Survey on Transfer Learning. IEEE Transactions on Knowledge and Data Engineering, 22(10), 1345–1359. doi: 10.1109/tkde.2009.191

23. Buolamwini, J., & Gebru, T. (2018). Gender Shades: Intersectional Accuracy Disparities in Commercial Gender Classification. Retrieved from https://proceedings.mlr.press/v81/buolamwini18a/buolamwini18a.pdf


Article Metrics

Metrics Loading ...

Metrics powered by PLOS ALM

Refbacks

  • There are currently no refbacks.




Copyright (c) 2024 Kolapo Obanewa, Olumide Innocent Olope

Creative Commons License
This work is licensed under a Creative Commons Attribution 4.0 International License.