Abstract
Managing carbon emissions from the manufacturing sector is crucial for sustainable development, and effective identification of manufacturing land is key to achieving this goal. However, current methods for identifying urban manufacturing land remain inadequate. In this study, we employ a fine-tuned, pre-trained natural language processing model based on Bidirectional Encoder Representations from Transformers to classify points of interest data into manufacturing industry categories. This approach enables us to identify manufacturing land and allocate corresponding carbon emissions data to specific parcels. The global Moran’s Index and local Moran’s Index are applied to analyze the relationship between manufacturing concentration and carbon emission intensity. The results demonstrate that the fine-tuned model achieved an accuracy rate of 91.6% on the test set, successfully identifying 98.72% of the manufacturing land in the study area. The intensity of carbon emissions from manufacturing exhibits a significant positive spatial correlation, with urban areas characterized by high-high and low-low clustering of emissions. In rural areas, high-emission manufacturers tend to be co-located with low-emission enterprises. Within individual manufacturing sectors, most exhibit low-low clustering, suggesting a potential relationship between such clustering and lower carbon emissions. This study provides detailed spatial data for the management of carbon emissions in the manufacturing sector and addresses the gap in micro-scale research on the correlation between manufacturing concentration and carbon emissions.
Keywords
Introduction
Carbon dioxide (CO2) is a principal greenhouse gas and a significant contributor to global warming (Hofmann et al., 2006). The Intergovernmental Panel on Climate Change (IPCC) report emphasizes that, in order to limit global temperature rise to within 1.5°, net-zero CO2 emissions must be achieved by 2050, a critical target for sustainable development (IPCC, 2020). As the world’s largest CO2 emitter (BP, 2017), China has pledged to achieve carbon neutrality by 2060 (Mallapaty, 2020). Research indicates that industrial energy consumption accounts for 66.16% of China’s total energy use, with manufacturing making up 80% of that industrial energy consumption (Tang and Liu, 2016). Therefore, reducing carbon emissions in the manufacturing sector is a crucial element of China’s dual-carbon goal, which has been highlighted in several national-level policy documents.
Spatial policies, such as spatial planning, land use planning, and urban design, are key approaches for guiding urban low-carbon development. Optimizing urban form and land use structure can reduce energy consumption related to transportation (Long et al., 2013). In addition, the readjustment of urban industrial land can effectively reduce carbon emissions related to industry (Shu and Xiong, 2019). To provide sufficient spatial details to support the reduction of urban carbon emissions through spatial policies, high-resolution carbon emission maps need to be drawn, which can reveal the differences in both the total amount and intensity of carbon emissions across various regions within a city; this information can be used to assess the potential for emission reduction in different areas and optimize land use layout (Wang et al., 2022).
On the other hand, the concentration of manufacturing activities is believed to have varying impacts on carbon emissions. However, there is no consensus in academia on whether this impact is positive or negative (Shen and Peng, 2021). As different industries have distinct energy use characteristics, it is crucial to differentiate their industrial structures when discussing the impact of industrial or manufacturing concentrations on carbon emissions (Zeng and Zhao, 2009). However, there is still a lack of experience in this area because it is impossible to measure the concentration characteristics of different manufacturing sectors within the city.
In light of this, we aim to address two key issues. First, we seek to map the carbon emissions of urban manufacturing sectors, which will reveal the emission levels and spatial distribution across different manufacturing categories. Second, we will analyze the relationship between manufacturing concentration and carbon emission levels within the city, providing crucial support for optimizing the layout of industrial land.
Literature review
Research on carbon emission maps
Current research has made many efforts in the field of manufacturing carbon emissions, concerning factors affecting them (Liu et al., 2022), the role of policies such as environmental regulations on manufacturing (Wang et al., 2018), and the potential for emission reduction in manufacturing (Yan and Fang, 2015). However, these studies usually focus on the national or provincial level, or treat urban manufacturing as a separate sector, and lack sufficient spatial details to support the management of manufacturing carbon emissions on the urban scale.
Drawing high-resolution carbon emission maps is very important for urban low-carbon management (Wang et al., 2014). Depending on the spatial carrying objects, these maps can include representations of land use (Zhang et al., 2023), buildings (Zhang et al., 2022), and abstract vector elements (Liu et al., 2020). Theoretically, carbon emission maps can be drawn through top-down and bottom-up approaches. The top-down approach involves calculating carbon emissions at the urban level, and then allocating the calculated results to parcels or building levels, based on specific weighting factors. For example, the total carbon emissions are calculated based on the city-level energy balance sheet and the carbon emissions are allocated to agricultural, residential, industrial, and commercial sectors based on features of different lands, such as land area, population, building floor area, and industry category (Chuai and Feng, 2019). The bottom-up approach depends on detailed survey data, and the carbon emission map is obtained by establishing a spatial connection between the survey data and spatial elements. Although the bottom-up method has a higher data accuracy rate, it is difficult to obtain data due to the large number of buildings at the city scale, and carbon emission maps can be drawn only using methods such as prototype buildings; and then the carbon emissions are allocated to all buildings, based on the results of building type classifications (Zhang et al., 2022). Research generally uses a top-down approach to draw carbon emission maps.
However, due to research data limitations, the carbon emissions of the industrial sector are usually treated as a whole, which is extremely unfavorable to managing the emissions of the industrial sector, especially manufacturing. Carbon emission maps should be drawn for different sectors of manufacturing. The difficulty in achieving this goal lies in distinguishing manufacturing land from industrial land, as government statistical data in China’s urban planning system typically categorizes land solely as industrial.
Although manufacturing constitutes a significant portion of industrial activities, industrial land cannot be entirely considered manufacturing land, even when focusing only on the dominant function of land parcels and disregarding ancillary structures, such as storage and office buildings on industrial land. For instance, electricity and gas production, although not categorized as manufacturing in the national economic industry classification, still falls under industrial land use according to GB50137-2011 standards (Ministry of Housing and Urban-Rural Development of the People’s Republic of China and General Administration of Quality Supervision, 2011). Moreover, the new type of industrial land labeled as M0 within industrial land predominantly hosts industries and structures unrelated to manufacturing. Due to the inclusion of numerous land uses unrelated to manufacturing in industrial land classification, precise identification of manufacturing land is challenging, limiting its efficient management. Point of Interest (POI) data offer a promising solution to this issue.
Application of POI data in functional classification
POI data contain rich text information, which can effectively identify manufacturing categories. This is because the names of POIs belonging to the manufacturing industry consist of company names. According to the regulations of the China Enterprise Name Registration and Management, the name must include information about the industry to which the enterprise belongs. Therefore, mining industry information from POI texts can effectively address the challenge of identifying manufacturing land use. Currently, POI data are widely used for identifying building functions (Deng et al., 2022), land functions (Pan et al., 2023), etc., generally by reclassifying the self-owned categories of POIs to identify the required functions, with only a few studies exploring their textual information. For example, TF-IDF is used to rank the keywords of POI to identify building functions (Chen et al., 2020) but cannot fully mine its semantic features.
On the other hand, directly reclassifying POIs will inevitably introduce some errors because the self-owned categories of POIs cannot completely correspond to the categories of buildings and land. Some complex POI categories cannot be simply reclassified into another category. For example, the category of companies and enterprises in Amap’s POI includes not only manufacturing but also industries, such as construction, wholesale, and retail. Due to this, existing research tends to delete it or directly classify it into a certain category (Deng et al., 2022; Zhang et al., 2023), which leads to a reduction in the number of POIs and consequently affects the accuracy of function identification. Bidirectional Encoder Representations from Transformers (BERT), as an important model in the field of natural language processing, can fully understand the context of the text. The fine-tuned pre-trained BERT model has been widely used for text classification (Sun et al., 2019). Therefore, it is feasible to use BERT to learn the naming rules of manufacturing industry POIs and apply them to text classification.
The relationship between manufacturing concentration and carbon emissions
Regarding the impact of industrial concentration on carbon emissions, current research lacks a consensus. Some scholars believe that the concentration of manufacturing industries will exacerbate environmental pollution (Cheng, 2016), perhaps because it leads to the expansion of corporate capacity and the behavior of some companies indulging in “free riding,” that is, unwilling to invest in environmental improvement. Other scholars believe that the concentration of manufacturing industries will reduce carbon emissions, as the knowledge spillover in such settings promotes the application of advanced technology, thereby lowering energy consumption (Dong et al., 2012). In addition, concentration can also share facilities for unified government management. Some feel that there is no simple linear relationship between manufacturing agglomeration and carbon emissions. According to the stage of industrial development, the degree of manufacturing concentration and environmental efficiency shows a U-shaped relationship (Shen and Peng, 2021). These studies usually treat manufacturing or industry as a whole. According to China’s industry classification, manufacturing includes 31 sub-industries, and industry includes mining, manufacturing, electricity heat gas, and water production and supply industries. As these industries have highly differing energy use characteristics, different industrial sectors need to be analyzed separately. Some scholars have used entropy indicators to analyze the correlation between different sectors and divided them into four major categories. They assessed the degree of industrial density and corporate concentration and examined its impact on SO2 emissions (Cai and Hu, 2022). However, research at the sub-city scale remains insufficient. Measuring overall manufacturing aggregation at the city level cannot capture the spatial differences within the city or the nuances of industry clustering.
In summary, there is currently a lack of effective methods to accurately identify urban manufacturing land, which prevents the creation of high-resolution carbon emission maps for manufacturing. Furthermore, few studies have explored the impact of manufacturing concentration on carbon emissions at the intra-urban scale. To address these gaps, this study employs a fine-tuned pre-trained BERT model to classify POI data, thereby identifying manufacturing land. Based on this classification, a detailed carbon emission map is produced, and the impact of manufacturing concentration on carbon emission intensity is analyzed using global and local Moran indices.
Study area, data and methodology
Study area
Hefei is the capital city of Anhui Province, with a permanent population of 9.465 million in 2021. As shown in Figure 1, its core urban area consists of four districts (Baohe District, Yaohai District, Luyang District, and Shushan District), while the surrounding areas include four counties (Feidong County, Feixi County, Changfeng County, and Lujiang County) and one county-level city (Chaohu City). The urban center is located at the heart of the administrative boundaries, bordered by Chaohu Lake to the south and mountains to the north. In 2017, the National Development and Reform Commission designated Hefei as one of the third batch of low-carbon pilot cities and required it to develop a corresponding low-carbon urban development plan, making the implementation of low-carbon planning and management particularly important. Location of Hefei city.
In terms of industry, Hefei has four major economic and technological development zones: Hefei High-Tech Industrial Development Zone (HFHTZ), Hefei Economic and Technological Development Zone (ETDZ), Hefei Xinzhan High-Tech Industrial Development Zone (XZHTZ), and Anhui Chaohu Economic Development Zone (CHETDZ). ETDZ and HFHTZ are located in Shushan District, XZHTZ in Yaohai District, and CHETDZ in Chaohu City. The ETDZ comprises northern and southern zones. Manufacturing serves as a core industry in Hefei’s economy, accounting for 20% of the city’s GDP. Due to the pressing need for low-carbon management in Hefei’s significant manufacturing sector, this study selects the city as its case study subject.
Data
The data include POI data, industrial land data, energy consumption data, building outline data, building height data, and nighttime light data. Among them, POI data and some building outline and height data come from Amap. Amap (https://lbs.amap.com/) is a major map service provider in China, offering API interfaces that allow users to retrieve POI data. To supplement the building outline and height data in Amap data, the study additionally used the building outline dataset (Zhang et al., 2022) and building height dataset (Wu et al., 2023), from the National Tibetan Plateau/Third Pole Environment Data Center (https://data.tpdc.ac.cn) and Zenodo (https://www.zenodo.org). The industrial land data is sourced from the Hefei city government department and the energy consumption data from the Hefei Statistical Yearbook. The nighttime light data is sourced from the Luojia-1 nighttime remote sensing dataset, with a spatial resolution of approximately 130 m. We applied cubic interpolation to resample the nighttime light data, increasing its pixel resolution to 25 m to match the scale of the land parcels.
Methodology
As shown in Figure 2, our research methodology consists of three steps. The first step involves cleaning and processing the POI data using the spatial analysis module in ArcGIS (Supplemental Math 1). In the second step, we apply the BERT model to classify the POI data and identify manufacturing land use. In the final step, energy consumption is allocated to specific land parcels based on the identification results, using building floor area and nighttime lighting as weighting factors. Carbon emissions are then calculated based on the allocated energy consumption. Methodology.
BERT model
BERT, an advanced natural language processing model based on the Transformer architecture, is renowned for its robust semantic understanding capabilities. We selected the BERT model based on its powerful semantic understanding capabilities in natural language processing, making it particularly suitable for text classification tasks in POI data. Previous studies have successfully employed BERT to classify POIs, utilizing spatial attributes to identify urban green spaces (Cao et al., 2024) and land functions (Wang et al., 2023), thereby providing critical reference and validation for our research. Our objective is to accurately identify POIs related to manufacturing. BERT, by generating deep semantic vectors, effectively captures complex semantic features within the text, thereby assisting us in classifying manufacturing-related POIs into more specific subcategories.
In this study, we fine-tuned Huggingface’s BERT-base-Chinese model for a manufacturing classification task. Text data were tokenized using AutoTokenizer, with sequences capped at 512 tokens. The dataset was divided into training (70%), validation (10%), and test (20%) sets, applying Stratified Shuffle Split for balanced category distribution. The model was fine-tuned over five epochs with a learning rate of 1e-6, batch size of 20, and a 0.5 dropout rate to prevent overfitting. To streamline data processing, we created a custom PyTorch dataset class. BERT’s feature vectors were passed through a linear layer corresponding to the target categories training used the Adam optimizer with cross-entropy loss, and confusion matrices were computed during training and validation to monitor performance. To ensure reproducibility, we applied a fixed random seed and saved model states after each epoch. The final model’s accuracy was evaluated on the test set.
Carbon emission accounting and allocation
After classifying manufacturing POIs using BERT, we use the POIs to identify manufacturing parcels. We then apply the frequency density method (Supplemental Math 2) to count the POIs within each parcel and identify specific manufacturing industries. Subsequently, we allocate energy consumption across these parcels and calculate carbon emissions in three steps, as shown in Figure 3. First, we separate manufacturing energy consumption from the industrial energy consumption reported in statistical yearbooks. Then, we distribute this energy consumption to the parcels using building floor area and nighttime light intensity as allocation indicators. Finally, we estimate parcel-level carbon emissions via the emission factor method, as detailed below: Carbon emission calculation and allocation process.
In the municipal statistical yearbooks, industrial energy consumption encompasses manufacturing, mining, electricity and thermal power, and water production and supply (NBS Survey Office in Hefei, 2024). Therefore, to calculate the energy consumption of manufacturing, it is necessary to subtract the non-manufacturing sectors from the total industrial energy consumption. These statistics include only large enterprises (defined by the National Bureau of Statistics of China as companies with annual revenues exceeding 20 million yuan since 2011(National Bureau of statistics of the People’s Republic of China, nd)). On one hand, mining resources are not well-developed in Hefei city (among the original database of over 40,000 enterprise POIs, only six are related to the “mining” industry), resulting in a small number of mining enterprises; on the other hand, the industries of electricity, thermal power, and water production and supply are generally monopolized by large state-owned enterprises, leaving virtually no room for small and medium-sized enterprises. Given the typically large scale of enterprises in these sectors, it is reasonable to exclude the energy consumption of small and medium-sized enterprises in the statistics. Thus, the formula for calculating urban manufacturing energy consumption is as follows:
To obtain the sector-specific energy consumption within the manufacturing industry, the energy usage characteristics of large-scale industries are employed as weights to decompose the total energy consumption of the manufacturing sector, through the following formula:
The calculation of carbon emissions adopts the IPCC emission factor method (Eggleston et al., 2006). As energy consumption has been converted to standard coal, the carbon content of standard coal in the referenced literature is set to 0.67 (Zhou et al., 2019).To allocate carbon emission data to each manufacturing sector by industry-specific land parcels, the study uses building floor area and nighttime light intensity as allocation weights to disaggregate the carbon emissions down to the parcel level.
Building floor area, representing the total usable space of a structure, is an indicator of economic activity scale and is, therefore, used for carbon emission allocation. Previous studies have demonstrated a significant positive correlation between building floor area and carbon emission levels (Lai and Lu, 2019; Teng and Yin, 2023). However, to account for the potential presence of vacant buildings, nighttime light intensity is included as an additional allocation indicator. Nighttime light intensity, which is closely linked to energy consumption, is widely used in carbon emission estimations (Amaral et al., 2005). Although lighting energy consumption typically represents only a small fraction of a building’s total energy use, especially in industrial facilities where production equipment is the primary energy consumer, building floor area remains a more accurate proxy for total energy demand. This is because larger areas generally require more energy for heating, cooling, lighting, and other functions. Nonetheless, energy consumption can vary based on occupancy rates. To better capture actual building activity, nighttime light intensity is used as a supplementary measure. Consequently, the weights are set at 0.7 for building floor area and 0.3 for nighttime light intensity, providing a more precise representation of a building’s total energy consumption by incorporating both physical size and nighttime usage patterns. The calculation formula for the allocation weights is as follows:
Spatial analysis
To analyze the spatial heterogeneity of carbon emissions in the manufacturing industry, the study employs Moran’s Index to measure the spatial autocorrelation of manufacturing industry carbon emissions, which includes the global Moran’s Index and local Moran’s Index. The global Moran’s Index is used to analyze whether there is overall spatial autocorrelation in the data and the local Moran’s Index to measure the clustering types of data in specific local areas. The local Moran’s Index can be classified into four types based on different correlation statistics: high-high clustering (H-H), low-high clustering (L-H), low-low clustering (L-L), and high-low clustering (H-L), each representing different clustering patterns of carbon emission levels in the manufacturing sector (Supplemental Math 3).
Results
Results of POI classification and manufacturing land identification
The study used PyTorch to fine-tune a pre-trained BERT model. Figure 4 illustrates the accuracy curve over 100 epochs, and the model fine-tuned at epoch 51 was selected for classifying POI texts. The validation set accuracy of this model reached 92.8% while the accuracy on the test set was 91.6%. These results demonstrate the advantages of BERT in handling enterprise name data, showcasing its ability to comprehensively understand and effectively classify naming conventions. Therefore, using this model for POI text classification is deemed feasible. Combining some manually classified results, the study has a total of 17,180 classified POIs. Accuracy curve for 100 epochs.
Figure 5 shows the POI identification results, with the majority being mixed manufacturing areas. A considerable portion of land, however, was not identified, accounting for 11.6% of the total number. Nevertheless, in terms of area, the identified regions cover 98.27% of the total area. The average area of manufacturing land parcels is 31,842 m2, while uncategorized parcels average only 4750 m2. These smaller parcels, typically less than 100 m in length and width, are mainly located in suburban and rural areas. Omitting their carbon emissions is acceptable within the margin of error, as their energy consumption minimally impacts overall manufacturing energy use. The number of manufacturing parcels by different categories identified through POI.
The total manufacturing area in the study area (including mixed types) is 203.054 km2. The mixed type (involving two or more manufacturing types) has the highest proportion, accounting for 27.19%. The top two manufacturing types by area percentage are non-metallic mineral product manufacturing (6.31%) and the agricultural and sideline food processing industry (6.10%). These industries have relatively low entry barriers and lack significant raw material restrictions. These results are presented in Table S1 in the Supplemental Material, while Table S2 provides detailed explanations for each manufacturing category (Supplemental Table S1, S2).
Results of carbon emission accounting and allocation
Since carbon emissions are calculated based on specific manufacturing sectors, the emissions from mixed type manufacturing are allocated to the corresponding sectors. The total carbon emissions from the manufacturing industry in the study area reached 7.45 million tons in 2021. According to Table S3 in the Supplemental Material (Supplemental Table S3), the key sectors contributing to these emissions include the manufacturing of chemical raw materials and chemical products (2.42 million tons) and non-metallic mineral products manufacturing (2.22 million tons). The variation in total emissions across sectors is primarily due to differences in production scale. However, focusing solely on total emissions is insufficient to capture the distinct emission characteristics of each sector, so we measure emission intensity as emissions per unit area. Notably, chemical raw materials, chemical products manufacturing, and non-metallic mineral products manufacturing exhibit high-emission intensities. Although the former only occupies 3.27% of the total manufacturing area, it contributes 32.53% of the total emissions, while the latter, with 5.27% of the area, accounts for 29.76% of total emissions.
To visually represent the spatial distribution of carbon emissions, the emissions are allocated to various manufacturing zones, as shown in Figure S1(a) in the Supplementary Material (Supplemental Figure S1). Spatially, high-carbon emission areas are located mainly in the suburbs, attributed to the relocation of urban manufacturing industries. The overall carbon emission intensity in the manufacturing sector exhibits a pattern of high intensity in the southwest and northeast directions, converging around the major economic development zones and manufacturing enterprises in Hefei.
To facilitate urban low-carbon management, we map carbon emissions to the smallest administrative level in China, administrative villages and communities. Using hotspot analysis tools, we analyze the hotspots and coldspots of manufacturing carbon emissions and intensity, as depicted in Figures S1(b) and (c) in the Supplementary Material, respectively. Hotspot areas for total carbon emissions closely align with the distribution of development zones. Within the study area, there are seven carbon emission hotspot regions, with the two major ones located in the northern and western sections. The northern hotspot area includes the XZHTZ, while the western hotspot includes the HFHTZ and the southern part of the ETDZ. The distribution of carbon emission intensity, measured by emissions per unit of building floor area, is more scattered in hotspot areas. The study conducted statistical analyses on carbon emission quantities and intensities in various districts and counties, identifying highly significant hotspot regions (p ≤ .001), as shown in Table S4 in the Supplementary Material (Supplemental Table S4). These regions are concentrated in Shushan District, Feixi County, and Changfeng County, with a few in Yaohai District.
Spatial autocorrelation and aggregation patterns of manufacturing carbon emission intensity
The global Moran’s I was employed to statistically assess the carbon emission intensity per unit area, to explore the spatial heterogeneity of carbon emissions in the manufacturing industry. The Moran’s I value was found to be 0.039, with 999 random permutations resulting in a p-value of 0.001 and a Z-score of 20.59. Consequently, an overall positive spatial correlation can be inferred in carbon emission intensity across the study area, which is highly significant. Spatial positive correlation implies that regions with high-carbon emission intensity tend to be spatially clustered, and the same holds for regions with low-carbon emission intensity. However, the global Moran’s I cannot measure specific clustering patterns. The study employed local Moran’s I for further analysis, and the resulting LISA clustering patterns are depicted in Figure 6. LISA clustering results.
Proportions of different cluster types in urban and rural areas.
L-L and H-H clustering patterns indicate that manufacturing enterprises with low and high-carbon emission intensities, respectively, tend to agglomerate in urban areas. In contrast, H-L clustering reveals that in rural areas, high-carbon emission enterprises are more likely to co-locate with low-carbon emission enterprises. These patterns reflect significant differences in the spatial concentration of manufacturing industries between urban and rural regions. In urban areas, the emergence of L-L and H-H clusters may be linked to industrial parks that concentrate enterprises from similar sectors, which often share comparable energy usage characteristics. In rural areas, however, the prevalence of H-L clusters suggests a distinct clustering dynamic, one unlikely to result from random distribution of manufacturing enterprises, as the proportion of L-H clusters is notably low. This phenomenon may be associated with urban industrial upgrading, where some high-emission enterprises are relocated to rural regions, attracting low-emission enterprises through supply chain relationships and the advantages of logistics and infrastructure. These causal relationships require further research and validation.
The study conducted separate global Moran’s I tests for 31 individual manufacturing sectors and mixed manufacturing to assess the spatial autocorrelation within a specific manufacturing sector based on the results summarized in Table S5 in the Supplementary Material (Supplemental Table S5). Most manufacturing industries exhibit significant spatial positive autocorrelation. Industries with stronger spatial autocorrelation include Leather, Fur, Feather, and their Products and Footwear Manufacturing; Textile Industry; Metal Products, Machinery, and Equipment Repair Industry; and the Textile, Clothing, and Apparel Industry. This suggests that these sectors may experience clusters of both high-carbon and low-carbon emissions. In contrast, industries with lower Moran’s I values, such as Mixed Types and the Metal Products Industry, tend to show an approximately random spatial distribution. The study further analyzed the specific clustering patterns for these sectors, as presented in Table S6 in the Supplementary Material ( Supplemental Table S6). In industries with statistical significance, the clustering pattern is predominantly L-L in most manufacturing sectors, while a few sectors, such as the Textile Industry and the Cultural, Educational, Artistic, and Sporting Goods Manufacturing Industry, exhibit a dominant H-L pattern. The prevalence of the L-L pattern likely reflects the success of environmental policies and the adoption of energy-saving technologies, as it indicates that firms with lower carbon emissions tend to cluster together within the same industry. In contrast, the H-L pattern may suggest uneven implementation of environmental policies and regulations, particularly in rural areas where enforcement tends to be less stringent.
Conclusions and discussion
Conclusions
(1) Our research demonstrates that fine-tuning pre-trained BERT models exhibit high adaptability to the naming conventions of manufacturing POI, achieving an accuracy of 91.6% on the test set. Compared to reclassifying or deleting POIs, leveraging BERT to fully exploit the textual information of POIs enables the identification of a large portion of manufacturing land, approximately 98.27% of the total area, using a distance-weighted frequency density method.
Concerning building and land use function recognition, the uneven distribution of POI data poses a significant challenge to the effectiveness of function recognition. Previous studies have addressed this issue by expanding data sources, integrating POI data with image or geospatial data to improve classification coverage (Deng et al., 2022). However, during the process of POI reclassification, directly removing difficult-to-classify POIs leads to a loss of valuable data. Our study shows that leveraging the textual information of POIs alone can effectively identify manufacturing land.
(2) The research indicates significant spatial heterogeneity in urban manufacturing carbon emissions, revealing a correlation between the intensity of carbon emissions from manufacturing and the spatial distribution of manufacturing land, with noticeable disparities between urban and rural areas. Most manufacturing industries exhibit significant spatial positive correlation, with the L-L clustering pattern being the predominant mode of concentration across various industries. This suggests that there is a certain correlation between the clustering of manufacturing industries and the reduction of carbon emissions.
In urban areas, L-L and H-H clustering are the dominant patterns, possibly due to industrial parks’ control over the nature of key industries, such as chemical industrial parks and circular economy parks. In contrast, in rural areas, the predominant clustering pattern is H-L, indicating that high-carbon-emitting enterprises tend to cluster with low-emission enterprises. This could be attributed to the relocation of some energy-intensive manufacturing industries from urban areas, which, by sharing industrial chains and infrastructure, have attracted a group of upstream and downstream companies. However, this hypothesis requires further research and validation.
From an intra-industry perspective within manufacturing, L-L clustering is the predominant clustering pattern for industries exhibiting spatial positive correlation. The clustering of these industries correlates with reductions in carbon emissions. It is worth noting the case of mixed manufacturing industries, where prior research has shown that increasing related diversity reduces SO2 emissions while increasing unrelated diversity has the opposite effect (Cai and Hu, 2022). Our study reveals similar effects on carbon emissions, without distinguishing specific types of mixed manufacturing industries.
Research discussions
Policy recommendations
Although some manufacturing sectors exhibit different clustering patterns, the predominant clustering model for most industries remains the L-L type. This is particularly evident in industries such as Leather, Fur, Feather, and their Products and Footwear Manufacturing, as well as the Metal Products, Machinery, and Equipment Repair sectors, where the L-L model dominates. This suggests that encouraging L-L type clustering of manufacturing, whether in urban or rural areas, could effectively reduce carbon emissions, potentially due to factors like infrastructure sharing and technological spillovers.
However, policies must still differentiate between specific manufacturing sectors. For instance, industries like textiles and chemical fiber manufacturing predominantly display H-L clustering patterns in rural areas, where high-emission enterprises tend to cluster with low-emission ones. It is worth noting that, while the textile industry is primarily H-L dominant, a significant proportion of it also follows the L-L pattern, indicating substantial potential for emissions reductions.
Similar trends are observed in sectors such as textile apparel and accessories manufacturing, and printing and reproduction of recorded media, where H-L clustering is predominant in urban areas. These patterns may suggest a form of “free-riding” behavior within supply chains, where some firms may rely on the lower emissions of others to offset their own higher emissions. To address this, governments should focus on mitigating such behavior by establishing more sophisticated energy consumption evaluation systems to assess firms’ energy efficiency. This would enable the implementation of differentiated incentive policies that encourage increased investment in environmental sustainability.
Discussion on results uncertainty, model limitation, and future concerns
This study incorporates building floor area and nighttime light as weighting factors to allocate carbon emissions to specific parcels. Among them, nighttime light effectively reflects the actual operational conditions of the manufacturing industry, with its weighting based on its direct correlation with electricity consumption (Amaral et al., 2005). According to the Hefei Statistical Yearbook 2022, approximately 27% of the energy consumption in the manufacturing sector is attributed to electricity. Although the inclusion of nighttime light has a significant impact on parcel-level carbon emission estimates, the dual-weighting approach, combining building floor area and nighttime light, demonstrates greater scientific rigor and applicability compared to using building floor area alone. Moreover, previous study has confirmed the reliability and effectiveness of nighttime light data in estimating industrial carbon emissions (Wei et al., 2024).
Nevertheless, the diversity of weighting factors in the current model still has room for improvement. Therefore, future research will explore the incorporation of additional weighting variables that reflect the dynamic characteristics of the manufacturing industry to further enhance the robustness and precision of the model. Although the weight settings have been carefully considered, the weight proportions of building floor area and nighttime light intensity will still have some impact on the results of carbon emission allocation.
This study employed the BERT model for the classification of POI textual data. However, the scope of textual data in urban analysis extends far beyond POIs, encompassing other forms of social perception data such as social media texts, which hold significant application potential. For instance, some scholars have explored leveraging social media images for building function classification (Hoffmann et al., 2023), indicating that the potential of social media textual data in this field remains underexplored and warrants further research.
At the same time, this study only verified the types of spatial autocorrelation and did not provide causal inferences between manufacturing clustering and carbon emissions. In future research, we plan to further validate this relationship using spatial causal inference models.
Supplemental Material
Supplemental Material - Using a machine learning framework for natural language processing to create a high-resolution carbon emission map for urban manufacturing
Supplemental Material for Using a machine learning framework for natural language processing to create a high-resolution carbon emission map for urban manufacturing by Tinyu Wang, Fengying Yan, Jian Ma and Xiaoping Zhang, Liang Dong in Environment and Planning B: Urban Analytics and City Science
Footnotes
Author Contributions
TY.W: Conceptualization, Methodology, Data curation, Formal analysis, Investigation, Writing- original draft. FY.Y: Methodology, Supervision, Resources, and Writing - review & editing. J.M: Writing - review & editing. XP.Z: Writing - review & editing. L.D: Methodology, Supervision, Resources, and Writing - review & editing.
Declaration of conflicting interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author(s) disclosed receipt of the following financial support for the research, authorship, and/or publication of this article: This work was supported by the Special Funds of the National Natural Science Foundation of China [grant number 42341207] and National Science Foundation of China, Key Project [grant number 52338002]. The last author also thanks to: National Natural Science Foundation, China (NSFC), and the Dutch Research Council (NWO) (NSFC-NWO, NSFC: 72061137071; NWO: 482.19.608). Environment and Conservation Fund, Hong Kong SAR (ECF 88/2022).
Data availability statement
Due to restrictions on the disclosure of some original data, we are unable to provide the final dataset for this study. However, we have made the code used in the study publicly available. The code has been made publicly available on Zenodo. Researchers in need can access it by submitting an application. (URL: https://doi.org/10.5281/zenodo.13911963). The original data used in the study is listed below, and interested parties can access it through the following methods: The industrial land use data was provided by the Hefei Municipal Government and cannot be publicly disclosed due to restrictions on data-sharing permissions. The building footprint data (1) is sourced from the AMaps Open Platform and obtained via its API. We do not have the rights to share this data publicly. Researchers interested in accessing this data can apply for an API key directly from AMaps. URL: https://lbs.amap.com/. The building footprint data (2), primarily covering suburban and rural areas, is sourced from the research conducted by Zhang Z et al. While we do not have the authority to share this data directly, it is stored in the National Tibetan Plateau/Third Pole Environment Data Center and can be downloaded directly from the platform by interested researchers. URL: https://data.tpdc.ac.cn (Zhang et al., 2022). The building height data is sourced from the study conducted by Wu WB et al. and is stored on Zenodo. We do not have the rights to share this data publicly; researchers interested in accessing it can download it directly from Zenodo. URL: (https://zenodo.org/record/7064268#.YxtVAuxBz0p (Wu et al., 2023). The energy statistics data is sourced from the Hefei Statistical Yearbook 2023 and obtained from the Hefei Bureau of Statistics website. URL: https://tjj.hefei.gov.cn/tjnj/2023nj/index.html (Hefei Municipal Bureau of Statistics, 2024).
Supplemental Material
Supplemental material for this article is available online.
References
Supplementary Material
Please find the following supplemental material available below.
For Open Access articles published under a Creative Commons License, all supplemental material carries the same license as the article it is associated with.
For non-Open Access articles published, all supplemental material carries a non-exclusive license, and permission requests for re-use of supplemental material or any part of supplemental material shall be sent directly to the copyright owner as specified in the copyright notice associated with the article.
