A Data Management Scheme for Micro-Level Modular Computation-Intensive Programs in Big Data Platforms

Chakroborti, Debasish; Roy, Banani; Mondal, Amit; Mostaeen, Golam; Roy, Chanchal K.; Schneider, Kevin A.; Deters, Ralph

doi:10.1007/978-3-030-32587-9_9

Debasish Chakroborti ORCID: orcid.org/0000-0002-1597-8162⁵,
Banani Roy⁵,
Amit Mondal⁵,
Golam Mostaeen⁵,
Chanchal K. Roy⁵,
Kevin A. Schneider⁵ &
…
Ralph Deters⁵

Part of the book series: Studies in Big Data ((SBD,volume 65))

1039 Accesses
1 Citations

Abstract

Big Data analytics or systems developed with parallel distributed processing frameworks (e.g., Hadoop and Spark) are becoming popular for finding important insights from a huge amount of heterogeneous data (e.g., image, text, and sensor data). These systems offer a wide range of tools and connect them to form workflows for processing Big Data. Independent schemes from different studies for managing programs and data of workflows have been already proposed by many researchers and most of the systems have been presented with data or metadata management. However, to the best of our knowledge, no study particularly discusses the performance implications of utilizing intermediate states of data and programs generated at various execution steps of a workflow in distributed platforms. In order to address the shortcomings, we propose a scheme of Big Data management for micro-level modular computation-intensive programs in a Spark and Hadoop-based platform. In this paper, we investigate whether management of the intermediate states can speed up the execution of an image processing pipeline consisting of various image processing tools/APIs in Hadoop Distributed File System (HDFS) while ensuring appropriate reusability and error monitoring. From our experiments, we obtained prominent results, e.g., we have reported that with the intermediate data management, we can gain up to 87% computation time for an image processing job.

This is a preview of subscription content, log in via an institution to check access.

Access this chapter

Log in via an institution

Chapter: USD 29.95; Price excludes VAT (USA)

eBook: USD 149.00; Price excludes VAT (USA)

Softcover Book: USD 199.99; Price excludes VAT (USA)

Hardcover Book: USD 199.99; Price excludes VAT (USA)

Tax calculation will be finalised at checkout

Purchases are for personal use only

Institutional subscriptions

References

Becker, T., Cavanillas, J., Curry, E., & Wahlster, W. (2016). Big Data Usage, New Horizons for a Data-Driven Economy. Cham: Springer.
Google Scholar
Bezzo, N., Park, J., King, A., Gebhard, P., Ivanov, R., & Lee, I. (2014). Demo abstract: ROSLab A modular programming environment for robotic applications. In ACM/IEEE International Conference on Cyber-Physical Systems (ICCPS), Berlin (pp. 214–214).
Google Scholar
Blomer, J. (2015). Experiences on file systems: Which is the best file system for you? Journal of Physics: Conference Series, 664(4), 042004.
Google Scholar
Coppens, F., Wuyts, N., Inz, D., & Dhondt, S. (2017). Unlocking the potential of plant phenotyping data through integration and data-driven approaches. Current Opinion in Systems Biology, 4, 58–63.
Article Google Scholar
Depardon, B., Mahec, G. L., & Seguin, C. (2013). Analysis of Six Distributed File Systems [Research Report]. pp. 44, hal-00789086.
Google Scholar
Desprez, F., & Dutot, P. F. (2016). Euro-Par 2016: Parallel processing workshops. In Euro-Par 2016 International Workshops, Grenoble, August 24–26, Lecture Notes in Computer Science.
MATH Google Scholar
Donvito, G., Marzulli, G., & Diacono, D. (2014). Testing of several distributed file-systems (HDFS, Ceph and GlusterFS) for supporting the HEP experiments analysis. Journal of Physics: Conference Series, 513(4), 04.
Google Scholar
Han, Z., & Hong, M. (2017, 27 April). Signal Processing and Networking for Big Data Applications. Cambridge: Cambridge University Press.
Google Scholar
Heit, J., Liu, J., & Shah, M. (2016). An architecture for the deployment of statistical models for the big data era. In IEEE International Conference on Big Data, Washington, DC (pp. 1377–1384).
Google Scholar
Kaseb, A. S., Mohan, A., & Lu, Y. H. (2015). Cloud resource management for image and video analysis of big data from network cameras. In International Conference on Cloud Computing and Big Data (CCBD), Shanghai (pp. 287–294).
Google Scholar
Kim, M., Choi, J., & Yoon, J. (2015). Development of the big data management system on national virtual power plant. In 10th International Conference on P2P, Parallel, Grid, Cloud and Internet Computing (3PGCIC), Krakow (pp. 100–107).
Google Scholar
Li, B., He, Y., & Xu, K. (2011). Distributed metadata management scheme in cloud computing. In 6th International Conference on Pervasive Computing and Applications, Port Elizabeth (pp. 32–38).
Google Scholar
Luyen, L. N., Tireau, A., Venkatesan, A., Neveu, P., & Larmande, P. (2016). Development of a knowledge system for Big Data: Case study to plant phenotyping data. In Proceedings of the 6th International Conference on Web Intelligence. New York: ACM, Article 27, 9 pages.
Google Scholar
Minervini, M., Scharr, H., & Tsaftaris, S. (2015). Image analysis: The new bottleneck in plant phenotyping [applications corner]. IEEE Signal Processing Magazine, 32(4), 126–131.
Article Google Scholar
Minervini, M., & Tsaftaris, S. A. (2013). Application-aware image compression for low cost and distributed plant phenotyping. In 18th International Conference on Digital Signal Processing (DSP), Fira (pp. 1–6).
Google Scholar
Mistrik, I., & Bahsoon, R. (2017). Software Architecture for Big Data and the Cloud. ISBN 9780128054673, Jun 12.
Google Scholar
Mondal, A. K., Roy, B., Roy, C. K., & Schneider, K. A. (2018). Micro-level Modularity of Computation-intensive Programs in Big Data Platforms: A Case Study with Image Data, Technical Report, University of Saskatchewan.
Google Scholar
Pineda-Morales, L., Costan, A., & Antoniu, G. (2015). Towards multi-site metadata management for geographically distributed cloud workflows. In IEEE International Conference on Cluster Computing, Chicago, IL (pp. 294–303).
Google Scholar
Pineda-Morales, L., Liu, J., Costan, A., Pacitti, E., Antoniu, G., Valduriez, P., et al. (2016). Managing hot metadata for scientific workflows on multisite clouds. In IEEE International Conference on Big Data (Big Data), Washington, DC (pp. 390–397).
Google Scholar
Prasad, S. K., et al. (2017). Parallel processing over spatial-temporal datasets from geo, bio, climate and social science communities: A research roadmap. In IEEE International Congress on Big Data (BigData Congress), Honolulu, HI (pp. 232–250).
Google Scholar
Roy, B., Mondal, A. K., Roy, C. K., Schneider, K. A., & Wazed, K. (2017). Towards a reference architecture for cloud-based plant genotyping and phenotyping analysis frameworks. In IEEE International Conference on Software Architecture (ICSA), Gothenburg (pp. 41–50).
Google Scholar
Singh, A., Ganapathysubramanian, B., Singh, A. K., & Sarkar, S. (2016). Machine learning for high-throughput stress phenotyping in plants. Trends in Plant Science, 21(2), 110–124. ISSN 1360-1385.
Google Scholar
Skidmore, E., Kim, S., Kuchimanchi, S., Singaram, S., Merchant, N., & Stanzione, D. (2011). iPlant atmosphere: A gateway to cloud infrastructure for the plant sciences. In Proceedings of the 2011 ACM Workshop on Gateway Computing Environments (pp. 59–64). New York, NY: ACM.
Chapter Google Scholar
Smith, K., Seligman, L., Rosenthal, A., Kurcz, C., Greer, M., Macheret, C., et al. (2014). “Big Metadata”: The need for principled metadata management in big data ecosystems. In Proceedings of Workshop on Data analytics in the Cloud. New York, NY: ACM, Article 13, 4 pages.
Google Scholar
Sun, W., Wang, X., & Sun, X. (2012). Ac 2012-3155: Using modular programming strategy to practice computer programming: A case study. American Society for Engineering Education.
Google Scholar
Tudoran, R., Nicolae, B., & Brasche, G. (2017). Data multiverse: The uncertainty challenge of future big data analytics. In Semantic Keyword-Based Search on Structured Data Sources. Lecture Notes in Computer Science (Vol. 10151). Cham: Springer.
Google Scholar
Uti, A. M., Brand, M. V. D., Verhoeff, T. (2017). Exploration of modularity and reusability of domain-specific languages: An expression DSL in MetaMod. Computer Languages, Systems and Structures, 51, 48–70. ISSN 1477-8424.
Google Scholar
Walter, A., Liebisch, F., & Hund, A. (2015). Plant phenotyping: From bean weighing to image analysis. Plant Methods, 11(1), 14.
Article Google Scholar
Wang, F., Qiu, J., Yang, J., Dong, B., Li, X., & Li, Y. (2009). Hadoop high availability through metadata replication. In Proceedings of the First International Workshop on Cloud Data Management (pp. 37–44). New York: ACM.
Chapter Google Scholar
Wu, S. G., Bao, F. S., Xu, E. Y., Wang, Y. X., Chang, Y. F., & Shiang, C. L. (2007, December). A leaf recognition algorithm for plant classification using probabilistic neural network. In IEEE 7th International Symposium on Signal Processing and Information Technology, Cairo.
Google Scholar
Yang, X., Liu, S., Feng, K., Zhou, S., & Sun, X. H. (2016). Visualization and adaptive subsetting of earth science data in HDFS: A novel data analysis strategy with Hadoop and Spark. In IEEE International Conferences on Big Data and Cloud Computing (BDCloud), Atlanta, GA (pp. 89–96).
Google Scholar

Download references

Acknowledgement

This work is supported in part by the Canada First Research Excellence Fund (CFREF) under the Global Institute for Food Security (GIFS).

Author information

Authors and Affiliations

Department of Computer Science, University of Saskatchewan, Saskatoon, SK, Canada
Debasish Chakroborti, Banani Roy, Amit Mondal, Golam Mostaeen, Chanchal K. Roy, Kevin A. Schneider & Ralph Deters

Authors

Debasish Chakroborti
View author publications
You can also search for this author in PubMed Google Scholar
Banani Roy
View author publications
You can also search for this author in PubMed Google Scholar
Amit Mondal
View author publications
You can also search for this author in PubMed Google Scholar
Golam Mostaeen
View author publications
You can also search for this author in PubMed Google Scholar
Chanchal K. Roy
View author publications
You can also search for this author in PubMed Google Scholar
Kevin A. Schneider
View author publications
You can also search for this author in PubMed Google Scholar
Ralph Deters
View author publications
You can also search for this author in PubMed Google Scholar

Corresponding author

Correspondence to Debasish Chakroborti .

Editor information

Editors and Affiliations

Department of Computer Science, University of Calgary, Department of Computer Engineering Istanbul Medipol University Istanbul, Turkey, Calgary, AB, Canada
Reda Alhajj
Department of Electrical and Computer Engineering, University of Calgary, Calgary, AB, Canada
Mohammad Moshirpour
Department of Electrical and Computer Engineering, University of Calgary, Calgary, AB, Canada
Behrouz Far

Rights and permissions

Reprints and permissions

Copyright information

About this chapter

Cite this chapter

Chakroborti, D. et al. (2020). A Data Management Scheme for Micro-Level Modular Computation-Intensive Programs in Big Data Platforms. In: Alhajj, R., Moshirpour, M., Far, B. (eds) Data Management and Analysis. Studies in Big Data, vol 65. Springer, Cham. https://doi.org/10.1007/978-3-030-32587-9_9

Download citation

DOI: https://doi.org/10.1007/978-3-030-32587-9_9
Published: 21 December 2019
Publisher Name: Springer, Cham
Print ISBN: 978-3-030-32586-2
Online ISBN: 978-3-030-32587-9
eBook Packages: EngineeringEngineering (R0)

Publish with us

Policies and ethics