Abstract
Carsharing and taxi are both shared or public car services in urban transportation systems, which means they have a lot in common. In particular, the narrow price gap between carsharing and taxi or public transit in China makes them similar. Trip characteristics analysis and market segmentation are needed, to investigate why and when travelers choose carsharing rather than taxi. In this study, vehicle GPS data and operation order data of a round-trip carsharing system in Hangzhou, China, is used to obtain information on 13,338 valid trips. The trips are divided into three groups based on the travel cost comparison with taxi. Then, an artificial neural network model is developed to analyze group characteristics. The trip characteristics concluded from the model, and analysis on typical service stations, reveal that carsharing has its price advantage on simple long-distance trips during off-peak hours, which is called the regular market. Carsharing’s extended market, in which travel cost is higher than by taxi, covers two types of trips: One involves short driving distance (around 20 km) and long stopping time, and tends to occur in areas in which it is difficult to hail a taxi during peak hours; the other involves very short driving distance (around 10 km), more activity spots and very low travel cost (around 25 CNY). The results of this study can help carsharing operators to extend and adjust their business. Also, these results can contribute to help city administrators to reach better decisions.
In recent years, carsharing has become an important transportation trend in urban areas. (“Carsharing” in this study specifically refers to time-charging rental service and does not include other online ride-hailing services like Uber and Lyft.) With its rapid growth, carsharing in China tends to be a new mode of traveling and attracts users from other public transport modes. This results from the travelers’ attitude to vehicles, shifting from owning to renting a vehicle service: On the one hand, besides the parking problems in urban areas, it has become increasingly difficult and expensive to own a car; on the other hand, using a vehicle service has become more efficient, according to the experience in North America ( 1 ). In addition, users can often get discounts when they use a carsharing service.
However, carsharing is not the only way to use transportation services. Public transport, taxi and car rental are among the alternative services available. Out of these modes, taxi service is the most similar to carsharing and both share many features:
Individuals use them because of their occasional need for a vehicle ( 2 ).
Individuals enjoy the use of a vehicle without owning one.
All costs are distributed across use, and there are no fixed costs ( 3 ).
Whether there is an available car is not guaranteed beforehand.
The duration of use would not be too long.
The travel is usually within the urban area.
Individuals enjoy space privacy while using the service.
A study pointed out that carsharing members use it to substitute for modes that most resemble driving ( 4 ), but there are few studies comparing taxi and carsharing services. It is quite necessary to conduct this comparison, especially in China. The cost of taking a taxi in China is cheaper than in many other countries ( 5 ), but the difficulty of hailing a taxi is becoming a severe problem in big cities. Fortunately, competition exists. It is important for carsharing operators, transportation engineers, as well as city managers to understand this phenomenon. Thus, what the regular market of carsharing services is, why travelers would prefer to choose carsharing outside of its regular market, and what the extended market of carsharing services is, are questions worth studying. This paper fills this gap as it develops an artificial neural network model to conduct an analysis of the impact factors, and it evaluates the regular market and the extended market. The results of this research can help carsharing service operators to recognize their market competitiveness compared with taxi services, and to develop more specific strategies.
In this study, the GPS data and operation order data of “Fun Carsharing,” the first operational electric carsharing system in China, will be used. “Fun Carsharing” is a round-trip carsharing system in Hangzhou. In this system, the rental vehicle must be returned to the same service station it was rented from. Therefore, one order’s GPS data can perfectly record one complete trip, which makes it possible to identify and analyze trip chains.
The paper is structured as follows. First it provides a review of studies on carsharing user behavior and the application of GPS data in the transportation field. Then the dataset used in this study and the data processing work are introduced. Concepts and methodology are described, followed by presentation of the results and discussion. The final section comprises the conclusions and future directions.
Literature Review
Behavior Studies on Carsharing
Recent studies on carsharing user behavior can be sorted into two groups according to their methods.
Many studies adopted questionnaire survey. At first, the potential users are simply described as social activists, environmental protectors, and innovators ( 6 ). Later studies gave more details. Surveys in Japan ( 7 ) and North America ( 8 , 9 ) found that short travel distance is a positive factor, whereas the conclusion is totally opposite in Shanghai, China ( 5 ), which implies that the condition of carsharing in China may be different from other countries. Other common positive factors are lower income, younger age, private car ownership, and environment consciousness. In addition, gender, duration, travel period, public transport policy, and convenience are also factors that affect ( 8 – 11 ).
Questionnaire survey is limited by its sample size. With the development of data acquisition and mining technology, people have started to analyze the pattern of use of carsharing and the classification of users based on big data ( 12 – 15 ). Detailed carsharing trip characteristics, such as distance, duration, and frequency, are concluded, and they differ in different cities.
These impact factors inspire the variable selecting in this paper. However, the main limitation of the above studies is that there are few studies on the comparison of carsharing with other travel modes, especially with taxi services, which have a great deal in common with carsharing.
Application of Vehicle GPS Data in the Transportation Field
Vehicle GPS data contains information about the time, position and speed of a vehicle. In recent years, GPS data has been successfully applied in the transportation field. Due to the privacy concerns of individuals, taxi GPS data is most common, but a few carsharing GPS data are also used.
Carsharing GPS data is used to analyze the conditions of service station ( 13 ) and trip chain characteristics ( 14 , 16 ). Certainly, taxi GPS data is also used to extract travel information ( 17 , 18 ), and reveal the operation conditions of taxi services ( 19 , 20 ). However, due to the large number of taxi vehicles, also human daily life mobility and activity information can be reflected by the passenger flow of taxis ( 21 ). In addition, taxi GPS data can be used to analyze, and provide advice about more efficient routes to taxi drivers ( 22 ). Moreover, anomaly detection and prediction are important applications of taxi GPS data ( 23 – 26 ).
In summary, though the number of carsharing vehicles limits the use of carsharing GPS data, the existing methods of taxi GPS data analysis inspire the methodology of this study. This paper combines operation order data and GPS data to obtain more details about the trips that operation order data cannot support, and also to overcome the limited sample size issue of the questionnaire survey method.
Dataset
Overview
The data used in this research is provided by the Hangzhou EVnet Co., the operator of the “Fun Carsharing” carsharing system in Hangzhou City. The operation data, as well as the vehicle GPS data, cover the period from December 2013 to June 2015, containing 31,446 rental records. The return interval of the GPS data is approximately 30 s, even though the car is not booked or moving, which makes it possible to recognize user activity spots. The two kinds of data can be linked by their common column “ORDERID” (Table 1).
Data Frame
Data Preprocessing
Errors often exist in the GPS data. To increase the accuracy and reliability of the analysis, data preprocessing is necessary. Ding ( 14 ) achieved some data cleaning jobs and cleared data records using the following principles:
The position of the first record of the rented state did not correspond to the service station
The average velocity between adjacent records is higher than 200 km/h ( 27 )
The change of average velocity does not satisfy the inequality ( 28 )
After that, the total number of valid orders is 28,553.
To reduce the impact of the GPS random error on travel distance, the GPS data is used to calculate the whole travel distance, and this is compared with the “DRIVING DISTANCE” in the order data. This can diminish the effects and potential problem caused by GPS drifting. Considering that different GPS devices still have 20% deviation in a round ( 29 ), then 20% is taken as the allowable error to select records with acceptable precision.
Moreover, to emphasize the service as an urban traffic mode for daily trips and to distinguish it from the conventional car rental mode, the selected orders are limited to a “RADIUS” within 50 km, which is just a bit further than the diameter of Hangzhou city, and with a “DURATION” of less than 24 hours.
Finally, a dataset with 13,338 orders is selected for the research.
Concepts and Methodology
Concepts and Coding
Before introducing the data processing and analysis process, it is necessary to define several new indices to label and describe the trip and to complete the following analysis. It is necessary to highlight that here the whole process of a rental is considered as a complete trip, which is different from the FHWA’s definition in the 2001 National Household Travel Survey ( 30 ). In addition to the terms D, PCS, R and ACTIVE mentioned above (Table 1), others are listed as follows, and a diagram helps to understand the indices better (Figure 1).

Trip with two activity spots.
Methodology
This paper (except for the discussion section) will include three main steps. First, group segmentation; data will be divided into three groups. Second, variable selection; large sample test will be used in this step. Third, modeling.
Step 1: Group Segmentation
Financial factors, for example travel cost, are mentioned frequently in existing research findings ( 31 – 33 ). Assuming that the users have considered the cost of both carsharing and taxi services beforehand, the trips are grouped by comparing the PCS and PT.
Considering that the pre-estimation of travel cost by people always slightly deviates from reality, a double-constraint tolerance range is set to judge whether taking a taxi is more economic. In this step, trips whose PT is 20% higher and costing 10 CNY more than PCS are assorted into Group I. Likewise, trips whose PT is 20% lower and costing 10 CNY less than PCS are assorted into Group III. The remaining trips are assorted into Group II. The result of group segmentation is demonstrated in the overview part of the next section.
Analysis of Group I and Group II will focus on the trip chain and member attribute characteristics. Analysis of Group III will place emphasis on the competitive advantage of the regular market of carsharing, based on the assumption of the rational-economic individual in economic theory.
Step 2: Large Sample Test
Before modeling, it is necessary to select variables suitable to incorporate into the model. A large sample test is used to test whether the sample has significant difference with the population. If so, it can be stated that this index affects the group’s choice behavior, to a certain degree.
In this step, eight indices will be tested, including D, R, T, PLST, ACTIVE, CENTER, PEAK and NIGHT. Taking the D of Group I as an example, the process of a large sample test is as follows:
The Null Hypothesis(H0) of all the tests is “there is no significant difference between sample and population.” Accordingly, the Alternative Hypothesis(H1) is “there is significant difference between sample and population.”
Test Statistic z of both type of variables is calculated as
where, n is the size of Group I,
The value of 0.01 is chosen as the Significance Level α.
If the absolute value of z is larger than the critical value
All the test results and discussions of these eight indices of the three groups are shown in the next section. Additionally, variables that have significant difference in at least two groups are capable of being predictors of the model.
Step 3: Artificial Neural Network Modeling
As there are high correlations among all the factors (Table 2), the common discrete choice model would not work well here. Therefore, the artificial neural network (ANN) is introduced.
Correlations Table
**Correlation is significant at the 0.01 level (2-tailed).
Correlation is significant at the 0.05 level (2-tailed).
The ANN model finds a hyperplane in multidimensional data space to split the data. It helps to determine differences among data groups. There are advantages in adopting ANN. First, ANN can deal with relatively large numbers of predictors and predictors with high correlation among them. Next, for carsharing service operators, with their operation, the more data they obtain, the better model they gain. Finally, the ANN model can work well in this study, benefiting from the large size of the data; particularly since about 30% of the data will be used for testing, not training, and if there is not enough data, the model would not be effective.
The main output of the ANN is two matrices of weight. One is from the input layer to the hidden layer; the other is from the hidden layer to the output layer. There is no direct weight from the input layer to the output layer. The hidden and output layers are calculated as
where,
The modeling work is conducted using the SPSS 22.0 software (IBM Corp., Armonk, NY, USA).
Results and Discussion
Overview of Group Segmentation and Large Sample Test
It is evident that the proportion of Group III (38%), whose PCS is clearly lower than PT, is the largest. However, an interesting finding is that the proportion of Group I (37%), whose PCS is obviously higher than PT, is just a little bit lower than that of Group III. This is probably due to the users’ strong preference for carsharing services, based on reasons associated with being conscious of convenience and the environment or their lifestyle and habits ( 11 ). Meanwhile, there is a moderate proportion of Group II (25%), whose PCS and PT are almost equal (Figure 2).

Group statistics and description.
According to the result of the large sample test, D, R, T, PLST ACTIVE and PEAK are incorporated into the model. D, R and PLST of the three groups are significantly different from the population, respectively; whereas T, ACTIVE and PEAK of two of the groups are significantly different from the population; and among the three groups, it does not matter whether it is night time or not. Also, the result reveals that trips in Group I occur during peak hours more often than in the general population, whereas trips in Group III are opposite to those in Group I. In addition, there are more active members in Group II, in contrast to Group III.
Result of ANN Modeling
In the final model, PCS is also selected because it still plays an important role. The correct percentage of the model falls sharply in the absence of PCS.Table 4 and Figure 3a carry the same information. The width and color of the lines in Figure 3a represent the values in Table 4. The correct percentage of Group I and Group III is 83% and 85% (Table 3), respectively, which are adequate. The correct percentage of the Group II is merely 43%, but the Group II-Group II pair is still the highest (Figure 3b and Table 3). The reason why the correct percentage of the Group II is around half of that of the other groups might be that Group II has both maximum and minimum thresholds in the group segmentation step, whereas the other groups only have one threshold. Moreover, the predictive ability of the model is appreciable, according to the cumulative gains chart (Figure 3c), in which a larger area under the curves indicates a better performance by the model. Thus, the result of the modeling is still acceptable.
Classification and Percent Correct

Showing (a) network diagram, (b) predicted-observed chart, and (c) cumulative gains chart.
Discussion
Regular Market: Analysis of Group III
The characteristics of this group of trips could reveal the competitive advantages of the regular market of carsharing services compared with those of the taxi services.
The data shown in Table 4 reveal that the most salient positive effect is from D and T1_1_3 = 1, whereas the most salient negative affect is from PCS, T1_1_3 = 3 and PEAK = 1. The fact that T1_1_3 = 1 has a positive effect whereas T1_1_3 = 3 has a negative effect implies that these kind of trips are simple, having less or even no activity spots. PEAK = 1 as a negative effect means that these trips tend to occur during off-peak hours.
Parameter (Weight) Estimates
The probability distribution plot (Figure 4a) reveal that most driving distances of Group III lie in the interval from 20 km to 60 km, and the total cost is less than 50 CNY. Considering that the radius of Hangzhou city is around 23 km, 20 to 60 km is long-distance travel for urban activity. Besides, the PLST of Group III is obviously lower, indicating that there is a relatively long driving time and short stop time, which correspond to “simple trip”.

Showing (a) probability distribution plot of continuous variables, (b) regular market of carsharing (sketch), (c) Oscar movie world service station surroundings land use classification, and (d) Sunshine pier service station surroundings land use classification.
In conclusion, the regular market of carsharing services compared with taxi services is those simple long-distance trips during off-peak hours (Figure 4b), and the long-distance characteristic matches the survey conducted by Wang et al. ( 5 ) in Shanghai. The regular market may involve trips with the purpose of delivering and picking somebody up.
Extended Market: Analysis of Group I and Group II
The characteristics of these groups could reveal how carsharing services get their extended market, and specifically, what these trips focus on.
The data for Group I shown in Table 4 reveal that the most salient positive effect is from PCS, PEAK = 1 and T1_1_3 = 2, whereas the most salient negative effect is from D and T1_1_3 = 1. As for Group II, the most salient positive effect is from PCS and T1_1_3 = 3, whereas the most salient negative effect is from D. These results imply that trips in Group I and Group II have more activity spots, but shorter travel distance. Compared with Group III, these trips are complex. Moreover, PEAK = 1 is a positive effect for Group I, implying that these trips tend to occur during off-peak hours, whereas Group II does not show this characteristic.
The combination of the probability distribution plots (Figure 4a), indicates that most of the driving distances of these two groups of trips lie at about 20 km; most of the values of Group II are even smaller. However, the travel cost of these trips shows different characteristics. Most of the costs of Group I are around 100 CNY. Group II has two peak values, namely about 25 CNY and over 100 CNY, and the former is significantly predominant. It can be stated that trips in the extended market tend to be of short driving distances and the travel costs are both low and high. The reason why short driving distances would result in high travel cost is that the PLST of Group I is larger. The charging standard considers both distance and time. Thus, although they are not driving, they still charged for stopping.
To examine the services from a deeper perspective, the service stations whose proportion of the trips predominantly belongs to Group I is analyzed. The proportion of trips belonging to Group I in the following service stations are over 60%. One is Oscar Movie World in the Xiacheng District, the other is Sunshine Pier in the Xihu District. A search on the online map reveal that they are both located in the central area. Oscar Movie World lies in the most economically prosperous area of Hangzhou, where land use is mainly commerce, including companies, restaurants, hotels and malls (Figure 4c). Land use classifications around Sunshine Pier are mainly education, finance and residential (Figure 4d). The land use around the service station implies that there might be high demand of travel due to entertainment and commuting. However, using carsharing services is more convenient and reliable, considering that it is extremely difficult to hail a taxi in Hangzhou, especially during peak hours ( 34 ). Thus, these can be the evidence to support the analysis of Group I and Group II.
Summing up all the above information, the extended market consists of two types of trips:
Type 1: short driving distance with long time stops at activity spots; travel during peak hours; might occur at areas where it is difficult to hail a taxi.
Type 2: short driving distance with more activity spots; low cost.
What needs to be emphasized is that the characteristics of Type 1 are highly correspondent with commuting.
Summary
The analysis performed on the real operation order data shows that travel cost is not the key factor of mode choice, at least not between carsharing services and taxi services. In contrast with taxi services, around two-thirds of carsharing trips are not covered by its regular market, implying that competition exists among these trips. Table 5 gives a brief review and summary of the analysis of the carsharing markets.
Regular and Extended Market of Carsharing
Conclusions and Future Directions
This paper analyzed three groups of trips and identified the regular and extended markets of carsharing services compared with taxi services, based on data mining and modeling results. The conclusions of this research can be summarized as follows:
The location of carsharing services is indeed different from that of taxi services in urban transportation systems because there is a specific regular market for carsharing services according to comparison with taxi services. This conclusion is consistent with Prettenthaler ( 3 ).
Since about two-thirds of the trips are outside the carsharing regular market, the cost difference is not the key, or at least not the only, impact factor. In addition, the non-price advantages of carsharing services do not exist in suburban trips as expected, implying that there is competition between carsharing and taxi services in some trips in the central area.
The carsharing market might expand due to its convenience and reliability, which coincides with the results reported by Prieto et al. ( 10 ) and Catalano et al. ( 31 ). For carsharing system operators, it is a good opportunity to expand and adjust their business. They should allocate more vehicles at service stations in areas where most trips fit the extended market. Additionally, they should notice that price battle is no longer the most effective strategy.
City administrators can recognize the weakness of the existing taxi system and realize that the development of carsharing is necessary. As the extended market grows, if they want to control the development of carsharing for the extended market, then there must be improvement in the taxi and public transit systems for commuters, since public transport policy is an important impact factor ( 10 ). Alternatively, they can implement beneficial policies for carsharing to encourage travelers to use it, and improve the transportation system service.
In this study, there are still issues to be solved. First, the exact purpose and detailed information of Type 2 of the extended market is not clear. Questionnaire survey directed at members who undertake this kind of travel is needed. Second, there are trips that probably involve using carsharing as a mode of commuting (Extended market, Type 1). Whether this use pattern should be encouraged or not is still worth discussing. Third, similar comparison should be extended to other carsharing systems to verify the result of this work.
Footnotes
Author Contributions
The authors confirm contribution to the paper as follows: study conception and design: Ying Hui, Cuiyi Liu; data collection: Ying Hui, Cuiyi Liu, Mengtao Ding; analysis and interpretation of results: Cuiyi Liu, Ying Hui; draft manuscript preparation: Cuiyi Liu. All authors reviewed the results and approved the final version of the manuscript.
The Standing Committee on Urban Transportation Data and Information Systems (ABJ30) peer-reviewed this paper (18-02383).
