Integration of Deep Learning, Geographic Information, and Agricultural Image Interpretation for Assisting Crop Planting Registration Applications

Integration of Deep Learning, Geographic Information, and Agricultural Image Interpretation for Assisting Crop Planting Registration Applications

Published: 2026.09.29
Accepted: 2026.09.21
3
Assistant Research Fellow
Agricultural Policy Research Center, ATRI, Taipei, Taiwan
Technical Manager
Interactive Digital Technologies Inc., IDT, Taipei, Taiwan
Technical Assistant Manager
Interactive Digital Technologies Inc., IDT, Taipei, Taiwan
Senior Systems Analyst
Ministry of Agriculture, Taipei, Taiwan

ABSTRACT

In recent years, agricultural development in Taiwan has been affected by multiple factors, including a declining birth rate, an aging farming population, industrial restructuring, rising labor costs, extreme weather events, and the increasing frequency of natural disasters. Under these circumstances, agricultural agencies rely on interpreting agricultural imagery, supplemented by on-site investigations, to support crop planting registration, cropland surveys, and agricultural disaster damage assessments. These efforts, however, remain heavily dependent on substantial human resources and therefore incur considerable labor costs. In response, the Department of Information Technology of Taiwan’s Ministry of Agriculture (MOA) used its geoportal training database service to export agricultural parcels belonging to the target crop class. The exported parcel data were then paired with corresponding imagery from the same observation period to construct a training dataset. A Mask R-CNN model was subsequently trained and integrated into ArcGIS Pro to facilitate agricultural image interpretation and data post-processing. Using rice as the target crop class, this study selected 18,781 agricultural parcels from multiple regions, including Taoyuan City, Hsinchu County, Miaoli County, Yunlin County, Chiayi County, Tainan City, and Hualien County, as the test dataset. After model-based image interpretation and data post-processing, most performance metrics, including precision, recall, and F1 score, ranged from 0.8 to 0.9. In addition to evaluating model performance, the post-processed prediction outputs generated complete spatial feature datasets and attribute tables for subsequent analysis. These datasets could be directly used to support field investigations, farmland management, and other related agricultural operations. The model also identified and located previously unidentified, sparsely distributed, or region-specific target features in the imagery. Overall, the proposed approach could reduce the labor and time required for manual surveys while facilitating more efficient agricultural management and the development of precision agriculture.

Keywords: Mask R-CNN, agricultural parcel, ArcGIS Pro, deep learning, crop planting registration

INTRODUCTION

In recent years, agricultural development in Taiwan has faced numerous structural challenges, including a shrinking labor force due to a declining birth rate, an aging farming population, the restructuring of traditional industries, rising labor costs, extreme weather events such as water shortages, and agricultural losses caused by natural disasters. To address these challenges, agricultural agencies often rely on aerial imagery, remote sensing data, farmers’ self-reporting, and labor-intensive field surveys conducted by agricultural investigators to register and verify crop cultivation.

Given increasingly challenging environmental conditions and persistent labor shortages, large-scale cropland surveys have become increasingly difficult and resource-intensive. The Department of Information Technology (DIT) of Taiwan’s Ministry of Agriculture (MOA) has been promoting a crop planting registration program that supports several objectives, including precision agricultural assessments, disaster response, and agricultural production and market analyses. Prior studies have demonstrated that deep learning can be applied to remote sensing imagery for land-cover and crop classification, indicating that artificial intelligence (AI) techniques can serve as effective supplementary tools for agricultural image interpretation (Kussul et al., 2017). Consequently, integrating agricultural imagery, deep learning techniques, and on-site field investigations to improve the efficiency of agricultural operations has become a key issue.

To this end, the DIT has collaborated with partner agencies over several years to collect nationwide imagery resources, delineate agricultural parcels, and develop agricultural services for the MOA Geoportal. The workflow begins with compiling aerial imagery from the relevant observation period and agricultural parcel feature datasets. Trained personnel review and edit these datasets based on the results of field investigations. The resulting model-ready datasets are then organized and uploaded to the MOA Geoportal training database service. Each uploaded dataset includes the crop type, the name of the corresponding imagery, the county or city, and other relevant attributes. Users can search for specific datasets through the web-based platform or access the Geoportal’s web layer services through ArcGIS Pro to retrieve data for model training and other downstream tasks.

In support of these MOA projects, this study adopted the Mask R-CNN architecture for model training. The trained model was integrated into a customized workflow developed using ArcGIS Pro ModelBuilder, after which a script was executed to automatically perform batch predictions on aerial imagery. The interpretation results were subsequently published through the MOA Geoportal agricultural platform service to support downstream agricultural image interpretation and field investigations.

By integrating aerial imagery, AI-based deep learning, geospatial information, and other resources, the proposed automated interpretation mechanism and its predicted feature outputs could support agricultural decision-making and policy analysis. The anticipated benefits include improving the efficiency of field survey data collection and reporting, accelerating the verification of cropland management information, and streamlining the processing of crop planting registration data. The object detection and instance segmentation capabilities of Mask R-CNN are particularly useful for accurately delineating parcel boundaries and identifying the locations of individual target features, thereby reducing the labor and time required for manual surveys. Within a modern agricultural management framework, this approach could provide a practical foundation for subsequent field investigations, agricultural management, and evidence-based decision-making.

PREPARATION OF DATASETS AND TOOL

To develop the agricultural image interpretation module, this study used rice as the target crop class. The training dataset was constructed using 1:5,000-scale DMC III aerial imagery provided by the Aerial Survey and Remote Sensing Branch of the Forestry and Nature Conservation Agency under Taiwan’s Ministry of Agriculture (MOA). The training parcel features were selected and exported from the MOA Geoportal training database service. Most of the DMC III imagery was acquired between April and June 2024, except for one image acquired in 2023. A total of 4,914 rice parcel features were included, covering 24 map-sheet areas across Taoyuan City, Hsinchu County, Miaoli County, Nantou County, Yunlin County, Chiayi County, Tainan City, Kaohsiung City, Pingtung County, Yilan County, Hualien County, and Taitung County.

This study selected Mask R-CNN as the deep learning model to identify and delineate rice parcels. Mask R-CNN is an instance segmentation model that combines object detection with pixel-level segmentation (He et al., 2017). It extends the Faster R-CNN architecture by adding a parallel branch for predicting object masks (Esri, n.d.-b). When detecting instances of the target crop class in an image, the model can accurately identify individual objects and delineate their boundaries at the pixel level. In previous agricultural applications, researchers used Mask R-CNN in ArcGIS Pro to delineate strawberry canopies and extract the corresponding spatial features, demonstrating the applicability of instance segmentation models to agricultural image interpretation (Abd-Elrahman et al., 2021).

The model uses a ResNet deep convolutional neural network as its backbone for feature extraction. The resulting feature maps capture image features at multiple levels, including edges, textures, semantic information, and spatial patterns (He et al., 2016).

The Region Proposal Network (RPN) then generates candidate object regions within an input image. Region of Interest Alignment (RoIAlign) then extracts features for each region of interest from the feature maps while preserving precise spatial alignment. Finally, the model generates three outputs: (1) a class label, which predicts the class of each detected object; (2) a bounding box, which defines and refines the spatial extent and location of each object; and (3) a mask, which provides pixel-level segmentation of the full spatial extent of each object (Figure 1).

By integrating ArcGIS Pro’s built-in tools, ModelBuilder, ArcPy scripts, and deep learning libraries, users can develop customized workflows to perform batch predictions with Mask R-CNN and post-process the resulting data across multiple target images.

IMAGE INTERPRETATION MODULE DEVELOPMENT AND EVALUATION

Through the MOA Geoportal training database service, this study obtained agricultural parcel feature datasets prepared by partner agencies using aerial imagery acquired during the corresponding observation periods (Figure 2). The platform allows users to filter training parcel features by county or city, crop type, aerial imagery map-sheet name, and other attributes. Users can then retrieve the selected features through the Catalog pane in ArcGIS Pro by connecting to the service (Figure 3). Finally, by consulting the feature attribute table, users can obtain a list of the imagery filenames associated with the selected training parcels.

Pre-processing of training data

Before exporting the training data, the agricultural parcel features were reviewed to ensure the target features were correctly labeled. They were also checked and edited against the corresponding aerial imagery. Common editing tasks included boundary adjustments, parcel splitting, parcel merging, and boundary redrawing (Figure 4). Feature labels were also verified, and atypical cases were reviewed to ensure consistent training data quality and improve model prediction accuracy.

Once editing was complete, the imagery filenames were recorded in the existing agricultural parcel attribute table and used to generate a list of DMC III aerial imagery files. Using this list, an ArcPy script batch-exported the training data in an Esri-compatible deep learning format (Figure 5).

Mask R-CNN model training

Before training the Mask R-CNN model, the training-data export parameters were configured using the default settings shown in Figure 6. In total, 24 aerial imagery map sheets were used to export the training data, generating 50,112 image chips for model training.

After preparing the training data, this study considered hardware resources, software configurations, and other technical requirements. An ArcPy script, deep learning libraries, and other required modules were then integrated within a Jupyter Notebook environment to train the Mask R-CNN model. After training, the model achieved an average precision (AP) score of 0.926 (Figure 7).

Post-processing of prediction results

Automated image interpretation was performed by integrating ArcGIS Pro’s built-in tools, ModelBuilder, and ArcPy scripts into a batch image prediction workflow. Using a sliding-window mechanism, the model generated predicted features for the target images based on its default inference settings (Figure 8).

Because the sliding-window mechanism divides each image into smaller, partially overlapping image chips, many predicted features did not capture the full spatial extent of individual parcels. To represent the predicted targets as complete parcels, the outputs were post-processed using ArcGIS Pro tools to merge fragmented features and perform spatial and statistical analyses (Figure 9). We analyzed 23 target images covering the test areas. These areas were located primarily in Taoyuan City, Hsinchu County, Miaoli County, Yunlin County, Chiayi County, Tainan City, and Hualien County. Figure 10 presents an overview of the rice parcel predictions and the corresponding attribute table for the test areas.

Model prediction results and accuracy assessment

To assess the performance of the Mask R-CNN model in identifying rice parcels, this study adopted several object detection metrics, including Intersection over Union (IoU), precision, recall, the F1 score, and average precision (AP). The precision–recall (P–R) curve was also used to evaluate prediction performance (Esri, n.d.-a), while the receiver operating characteristic (ROC) curve was included as a supplementary assessment measure (Fawcett, 2006).

The test dataset contained 18,781 parcel features that served as ground-truth rice parcels across 23 1:5,000-scale aerial imagery map sheets. The test areas covered Taoyuan City, Hsinchu County, Miaoli County, Yunlin County, Chiayi County, Tainan City, and Hualien County. Before model evaluation, the rice prediction outputs underwent additional post-processing, including removing fragmented features, reviewing and editing ground-truth data, and delineating areas of interest (AOIs). These procedures reduced potential sources of interference in the prediction and evaluation processes (Figure 11).

After prediction and post-processing, the model generated 20,475 predicted features across the test areas. Model performance was evaluated using an IoU threshold of at least 0.5, while the inference process applied ArcGIS Pro’s default high-confidence threshold of at least 0.9.

During performance evaluation, the model produced 16,892 true positives (TPs), 3,583 false positives (FPs), and 1,889 false negatives (FNs). These results corresponded to a precision of 0.825, a recall of 0.899, and an F1 score of 0.861. The interpolated average precision at an IoU threshold of 0.5 (AP@0.5) was 0.859 for predictions meeting the high-confidence threshold, while the area under the ROC curve (AUC) was 0.871 (Figure 12).

The P–R curve showed that precision remained relatively high until recall reached approximately 0.8, after which it gradually declined. This pattern indicated that the model maintained relatively strong detection performance across a broad range of recall values. The ROC curve also approached the upper-left region and yielded an AUC of 0.871, providing supplementary evidence that the model could discriminate between rice and non-rice parcel features.

 

CONCLUSION

This study integrated deep learning, ArcGIS Pro, the MOA Geoportal training database service, and related tools to develop an agricultural image interpretation module, perform data post-processing, and test and validate the model. In the test area, the model achieved a precision of 0.825, a recall of 0.899, and an F1 score of 0.861. Most of the other evaluation metrics also ranged between 0.8 and 0.9. The precision–recall (P–R) curve demonstrated relatively stable performance, while the receiver operating characteristic (ROC) curve, which was used as a supplementary evaluation measure, approached the upper-left corner, indicating satisfactory discrimination between the target and non-target classes.

After batch predictions were performed across the target area, the outputs were post-processed to directly represent spatial information, including the shape and distribution of the predicted features. The accompanying attribute tables contained the interpretation results, imagery sources, and other relevant attributes.

The interpretation results were also published through the MOA Geoportal agricultural platform service. For personnel engaged in field surveys, farmland management, and crop planting registration, these results could reduce the time, labor, and associated costs required for manual operations. When target features were sparsely and irregularly distributed throughout the imagery, the model’s object detection and localization capabilities could further accelerate manual inspection and improve operational efficiency.

By using existing imagery resources and exporting selected target features from the MOA Geoportal training database service, this study established a preliminary systematic workflow for generating training data. A Mask R-CNN model was trained to identify a single crop class, and GIS tools were integrated into automated workflows to support batch image prediction. The interpretation results provided detailed information on the shape and spatial distribution of crop features. This image interpretation approach could support subsequent agricultural decision-making, field surveys, and data reporting, thereby improve operational efficiency and reducing costs. It could also be extended to agricultural production management, disaster damage assessment, and other related applications.

REFERENCES

Abd-Elrahman, A., Britt, K., & Whitaker, V. (2021). A step-by-step guide for automated plant canopy delineation using deep learning: An example in strawberry using ArcGIS Pro software: FOR372/FR441, 9/2021. EDIS, 2021(5). https://doi.org/10.32473/edis-fr441-2021

Esri. (n.d.-a). How Compute Accuracy For Object Detection works. ArcGIS Pro documentation. Retrieved June 16, 2026, from https://doc.esri.com/en/arcgis-pro/latest/tool-reference/image-analyst/h...

Esri. (n.d.-b). How Mask R-CNN works? ArcGIS API for Python. Retrieved June 16, 2026, from https://developers.arcgis.com/python/latest/guide/how-maskrcnn-works/

Fawcett, T. (2006). An introduction to ROC analysis. Pattern Recognition Letters, 27(8), 861–874. https://doi.org/10.1016/j.patrec.2005.10.010

He, K., Gkioxari, G., Dollár, P., & Girshick, R. (2017). Mask R-CNN. In Proceedings of the IEEE International Conference on Computer Vision (pp. 2961–2969). IEEE. https://doi.org/10.1109/ICCV.2017.322

He, K., Zhang, X., Ren, S., & Sun, J. (2016). Deep residual learning for image recognition. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (pp. 770–778). IEEE. https://doi.org/10.1109/CVPR.2016.90

Kussul, N., Lavreniuk, M., Skakun, S., & Shelestov, A. (2017). Deep learning classification of land cover and crop types using remote sensing data. IEEE Geoscience and Remote Sensing Letters, 14(5), 778–782. https://doi.org/10.1109/LGRS.2017.2681128

Comment