Abstract
Identity theft targeting American citizens continues to rise. While prior research has linked this trend to data breaches and scams, a more recent development—check theft from U.S. Postal Service (USPS) mailboxes—may also be contributing to the problem. However, the absence of official data connecting check theft to identity theft limits our ability to fully assess this relationship. This paper investigates whether check theft activity observed on Telegram, a major platform for illicit trade, can predict officially reported identity theft incidents as recorded by the Financial Crimes Enforcement Network (FinCEN). By integrating open-source intelligence (OSINT), we were able to collect and analyze publicly available data and compare it with formal crime reports, to illustrate how combining alternative and official data sources can enhance our understanding of the online fraud ecosystem, particularly the factors driving identity theft.
Introduction
Financial exploitation, particularly through the theft of personal and financial information, continues to be a critical issue affecting millions of people globally. In the U.S. alone, financial losses from fraud have seen a dramatic rise. According to U.S. Government Accountability Office (GAO) between 2018 and 2022, the federal government lost between $233 Billion to $521 Billion dollars annually to fraud (U.S. Government Accountability Office, 2024). This sharp increase can be attributed to various factors, including the proliferation of digital technologies, the increasing sophistication of fraud techniques, and the ongoing vulnerabilities in traditional forms of communication, such as the postal system.
Fraudsters today employ a diverse range of methods to obtain personal and financial information, with the evolution of online fraud reflecting increased sophistication. While tactics leveraging advancements in technologies such as phishing, social engineering, and hacking have become common practices (Girish & Bhowmik, 2023). One traditional method which was resurrected in 2021 is mail theft. The U.S. Postal Inspection Service reported a significant 161% rise in mail theft complaints between 2020 and 2021, with a total of 299,020 incidents (Biegelman, 2009; U.S. Postal Service Office of Inspector General (USPS), 2021). Among these incidents, check theft alone accounted for losses amounting to $18 billion. This method of theft is especially alarming, as it allows criminals not only to exploit victims’ financial resources but also to sell stolen financial information on online illicit markets, where they can be used to commit additional fraudulent activities.
The growing use of encrypted messaging platforms, particularly Telegram (Shah et al., 2020), has exacerbated the challenge of addressing stolen check sales, as criminals increasingly rely on these platforms to conduct illicit transactions. Telegram has become a hub for various criminal activities, including the sale of stolen personal information, facilitated by its strong privacy features and large user base. Research has shown that illicit data markets on Telegram bring together thousands of users in unregulated environments, contributing to the proliferation of identity theft and financial fraud (Dehghanniri & Borrion, 2021; Holt & Lee, 2022).
Despite this growing problem, there has been a notable gap in research addressing stolen checks and their sale in online markets. While financial fraud is widely studied, the specific dynamics of stolen check sales in these illicit environments have received little scholarly attention. This paper seeks to examine whether check theft data predict identity theft in the USA. Specifically, we seek to determine whether the type of check stolen (personal or business check) is likely to impact on the probability of identity theft.
We answer these questions by conducting a detailed analysis of stolen checks using data scraped from 80 online illicit markets that advertised checks for sale between September 2021 and December 2022, and official SARS report data of stolen checks and stolen identities reports from Financial Crime Enforcement Network. Utilizing Optical Character Recognition (OCR) technology, we extracted and analyzed information from these stolen checks, allowing us to gain valuable insights into the characteristics of victims, the financial losses suffered by individuals and banks, and the overall scope of the issue for both businesses and individuals. By leveraging Open-source intelligence (OSINT) methods and combining them with sophisticated data extraction and analysis techniques, the study offers a comprehensive view of the extent and nature of check fraud victimization across online illicit markets.
Literature Review
Fraud and Identity Theft
Fraud and identity theft are deeply interconnected, with each crime often serving as a catalyst for the other. Identity theft, as defined by the Federal Trade Commission, involves the unauthorized acquisition and use of someone’s personal information for illicit purposes. Once fraudsters obtain access to sensitive information such as Social Security numbers, bank account details, or credit card numbers, they can commit a range of fraudulent activities, including unauthorized purchases, loan applications, and accessing financial accounts. The damage caused by these crimes is not limited to immediate financial losses; victims often suffer long-term harm to their credit scores and personal reputations, which can take years to repair.
Financial fraud, on the other hand, involves deliberately misrepresenting or omitting facts for the purpose of monetary gain (Garkava et al., 2024; Holt & Lee, 2022). These crimes often involve some degree of identity manipulation or impersonation, tying them closely to identity theft. Together, identity theft and financial fraud create a vicious cycle where one crime perpetuates the other, often leading to severe financial and personal consequences for the victims.
The methods used by criminals to obtain personal and financial information have evolved significantly over the past two decades. In the past, information was primarily acquired through methods such as dumpster diving, pickpocketing, and insider corruption. However, modern fraudsters increasingly rely on digital techniques, including phishing attacks and social engineering. One method that has persisted despite technological advances is mail theft, particularly the theft of checks. Mail theft has been a concern in the United States for over two centuries (Biegelman, 2009), but recent trends show a marked increase in theft targeting checks, which are often altered and fraudulently cashed.
As noted by Maimon (2022), fraudsters now focus on areas with high mail traffic, such as postal vehicles, collection boxes, and neighborhood delivery systems. These criminals frequently use stolen or duplicated keys to break into mailboxes and steal checks. Once obtained, these checks can be altered, such as by changing the payee’s name or amount, and then cashed fraudulently. Alternatively, the checks are sold on illicit online platforms, where they become part of a larger ecosystem of stolen personal information traded by cybercriminals.
Stolen Checks in Online Illicit Markets
Once fraudsters have possession of stolen checks, they either deposit the checks directly or sell them on illicit markets (Maimon, 2022). These online markets, operating on the darknet and encrypted messaging platforms like Telegram, have become key hubs for the sale of stolen personal and financial information (Garkava et al., 2024; Shah et al., 2020). Howell and Maimon (2022) describe this phenomenon as the “supply chain of purloined personal data.” These markets facilitate the sale of a wide array of stolen goods, including Social Security numbers, credit card details, and checks, contributing to the rise of identity theft and financial fraud (Howell et al., 2022; Ouellet et al., 2022).
Historically, these illicit markets were primarily hosted on the darknet, an obscure part of the internet that requires special software to access and is not indexed by conventional search engines. However, in recent years, cybercriminals have increasingly migrated to encrypted messaging platforms such as Telegram to conduct their illicit transactions (Howell & Maimon, 2022; Shah et al., 2020). Telegram is a cloud-based messaging application that provides users with end-to-end encryption, making it a popular choice for individuals seeking privacy and security. While Telegram hosts legitimate channels covering a wide range of topics, it has also become a hub for criminal activity, including the sale of stolen checks.
Telegram’s encryption features present a significant challenge for law enforcement, as they make it difficult to monitor or intercept communications between offenders. Despite these challenges, Telegram has become a valuable intelligence source for understanding trends in the sale of stolen checks and other personal data. As Pomerleau and Maimon (2022) explain, data extracted from Telegram channels can provide detailed insights into the prevalence of check theft, the characteristics of victims, and the impact of security measures designed to curb this activity. These data are critical for identifying emerging threats and developing strategies to combat financial fraud and identity theft.
While there is a substantial amount of literature discussing financial and check fraud, the majority of these studies conceptualize fraud and check fraud as either a white-collar crime or an act that involves some form of contact and cooperation between the offender and the victim. However, in recent years the offender confiscates checks from mailboxes and advertises them on online illicit markets. Other studies that examine online illicit markets examine the organizational structure of the markets and the mechanisms that enable their operations, and the use and impact of the commodities sold on those markets on other types of fraud and crimes (Aldridge & Décary-Hétu, 2016; Dehghanniri & Borrion 2021; Holt et al., 2015). Moreover, those studies are mainly focused on darknet markets and on small number of markets at the time. Consequently, these studies do not provide insights into the growing problem of stolen checks advertised online. Addressing this gap, this paper analyzes open-source cyber-intelligence data on checks advertised on online illicit markets obtained from Telegram to assess the extent of the issue and shed light on victim characteristics.
From Official Sources to Open-Source Intelligence
In the United States, four main organizations are responsible for collecting data on identity theft and check fraud: the FBI’s Internet Crime Complaint Center (IC3), the Federal Trade Commission (FTC), the National Crime Victimization Survey (NCVS), and Advanced Fraud Solutions (AFS). Each of these bodies relies primarily on self-reported data to estimate the prevalence of these crimes. As a result, the reported figures on cases numbers and financial losses vary significantly across sources and there might be significant under-reporting, making it difficult to assess the true scope of the problem.
For instance, the FBI’s IC3 reported 12,876 cases of identity theft in 2024, resulting in approximately $199.9 million in losses. In contrast, the FTC reported over 1.1 million identity theft cases during the same year, with estimated losses totaling $12.5 billion. These discrepancies highlight the limitations of existing data collection systems and the challenges in forming a reliable national estimate.
Over relying on self-reporting of financial fraud can be problematic for several reasons. First, research has consistently shown that individuals tend to underreport their victimization in surveys (Beals et al., 2017), often due to embarrassment, lack of awareness, or perceived futility in reporting. Second, there is a temporal lag between the theft of personal information and the discovery of its fraudulent use, leading to delays in reporting the incident to authorities. Third, financial crimes are often reported to non-law enforcement entities, which results in further inconsistencies in official statistics.
OSINT has become an increasingly valuable tool for supplementing official data on financial fraud, particularly given the limitations of traditional reporting systems (Ablon et al., 2014; Leukfeldt, 2017). OSINT refers to the practice of collecting and analyzing publicly available information (e.g., social media platforms, online forums, websites, news reports, academic publications, government databases, and illicit marketplaces on the surface, deep, or dark web) to generate actionable insights (Bazzell, 2023). Its appeal lies in its non-intrusive nature, cost-effectiveness, and real-time accessibility. For example, researchers studying cybercrime can monitor Telegram channels, Reddit threads, or darknet forums to identify emerging trends, actors, and methods in online fraud or identity theft (Ablon et al., 2014; Leukfeldt, 2017). However, while OSINT provides access to large volumes of data, its validity and reliability must be carefully assessed.
Conceptual Framework
In the context of check theft and identity misuse, it is essential to distinguish between check issuers and check recipients, as each faces distinct types of risk exposure with measurable financial consequences. Issuers, individuals, or businesses writing checks, sensitive financial information such as bank account and routing numbers, are being exposed. This data can be exploited by unauthorized parties to create counterfeit checks, initiate fraudulent withdrawals, or orchestrate full account takeovers. In a case of a stolen check the issuer may experience direct financial losses, depending on how quickly the fraud is detected and the protections offered by the bank. In particular, when the check is cashed by unauthorized party or when forged or altered check is successfully cashed, the issuer bears the immediate financial hit until a dispute is resolved, which can take weeks or event months.
Recipients, on the other hand, are at risk when their identity is used by fraudsters to illegally cash checks that were issued in their name—often with fake IDs or forged endorsements. When a check made out to a legitimate recipient is intercepted and cashed by an impostor, the rightful recipient may never receive the funds, while the check issuer still sees the money withdrawn from their account. In this scenario, both parties suffer: the issuer loses the full amount of the check, while the recipient may face delays in payment, loss of income, and reputational damage.
Although a recipient’s name alone may not enable full identity theft, it facilitates partial impersonation, especially when paired with other stolen credentials or falsified documents. This study adopts an operational definition of identity theft that includes the impersonation or unauthorized use of another individual’s financial identity, commonly through stolen or fraudulently negotiated check instruments. Therefore, we treat both issuers and recipients as financial victims, with check-related fraud often leading to substantial monetary losses, disrupted financial operations, and long-term impacts on trust in traditional payment mechanisms.
The Current Study
This study explores the intersection of check theft and identity theft by leveraging open-source intelligence (OSINT) data collected from Telegram, a platform increasingly used by actors involved in financial fraud. While official data sources such as those provided by the Financial Crimes Enforcement Network (FinCEN) offer important insights into reported financial crimes, they often lack granularity, timeliness, and context. To address these gaps, this research poses four guiding questions. First, we examine how check theft data gathered from Telegram corresponds with official reports of check theft submitted to FinCEN (Research Question 1). This comparison allows us to evaluate the alignment, or divergence, between underground market activity and formal reporting systems.
Second, we assess the extent to which Telegram-based OSINT can offer insights that go beyond what is captured in official data (Research Question 2). Given that many check theft incidents may go unreported or unclassified in public data sources, understanding the added value of OSINT is essential for building a more comprehensive fraud intelligence picture.
Third, we investigate whether trends in check theft observed on Telegram serve as a predictor for officially reported incidents of identity theft, and whether this predictive relationship has remained stable over time (Research Question 3). This analysis is intended to probe the temporal dynamics between these two types of financial crime and assess causality or correlation.
Finally, we explore whether certain categories of stolen-check personal, business, government, or other—are more likely to result in identity theft (Research Question 4). Identifying which types of checks are most commonly associated with downstream identity misuse can inform targeted prevention strategies by both public agencies and private institutions. Together, these questions aim to bridge the gap between formal crime reporting and emerging sources of cyber-fraud intelligence, offering a multidimensional perspective on the evolving threat landscape of identity theft.
Data and Methodology
Methods
To answer our research questions, we employed open-source intelligence (OSINT) methods for gathering data on fraudulent checks promoted in online illicit markets and retrieve FinCen SARS reports of suspicious check activities (Available at: https://www.fincen.gov/reports/sar-stats). OSINT refers to techniques that involve collecting publicly accessible data from open and online sources, including websites, social media platforms, forums, and dark web marketplaces (Bazzell, 2023). OSINT has increasingly become a cornerstone of cybercrime investigations, offering non-intrusive, high frequency access to intelligence sources previously untapped by traditional criminology (Yeboah-Ofor & Brimicombe, 2018). These methods have been widely used in cybersecurity research for identifying criminal behavior, financial fraud, and other illicit activities, as they provide valuable insights into underground economies and emerging fraud schemes (Ablon et al., 2014; Howell et al., 2023; Leukfeldt, 2017; Ouellet et al., 2022).
Data Collection and Procedures
Data Retrieval
The primary data for this study was collected online fraud illicit markets, particularly Telegram channels and groups advertising Personally Identifiable Information (PII), including stolen identities, credit cards details, compromised bank accounts, and other sensitive PII. These platforms provide a wealth of information including images of commodities, their pricing, and the methods used to commit fraud. The OSINT process began by identifying relevant marketplaces through keyword searches and snowballing techniques. The Telegram markets were identified through a purposive approach in which we actively sought for markets that advertised stolen checks, through keywords search (Check, cheque, Fuzzl) and snowball sampling from group or channel to another. This purposive approach ensured our dataset specifically captured the online illicit activity relevant to our research question.
To ensure comprehensive coverage, data was retrieved from 80 Telegram-based financial fraud illicit markets. This included scraping relevant marketplace, downloading images of PII commodities, and monitoring public discussions between sellers and buyers. Data collection occurred over 16 months, from September 2021 to December 2022 (See Figure 1).

Data collection flowchart from Telegram (September 2021–December 2022).
To ensure the accuracy and relevance of the data set, each image underwent a rigorous screening process. A team of graduate research assistants first screened the images to identify those containing checks, while removing any images related to other forms of financial fraud (e.g., compromised bank accounts, credit cards). After identifying relevant images, they were manually reviewed to confirm their pertinence and remove duplicates, ensuring that data was not redundant.
Once the relevant images were isolated, they were processed using Optical Character Recognition (OCR) technology to extract the textual information embedded within them. The OCR tool was configured to detect alphanumeric characters, ensuring the accurate extraction of key details such as bank name, check issuer information, check addressee, and check amount. The OCR output was then stored in a structured CSV format for subsequent analysis. After the OCR process, the extracted text was manually reviewed and corrected to ensure data accuracy.
These data collection methods adhered to ethical guidelines for researching illicit online activity. No personal information of victims was disclosed or exploited, and all findings were anonymized in the presentation of results. The use of OSINT also ensured that no interactions with cybercriminals took place, minimizing the risk of legal or ethical breaches in the course of data collection.
Additionally, to assess the validity of the OSINT data, we analyze official reports published by the Financial Crimes Enforcement Network (FinCEN). FinCEN, a bureau within the U.S. Department of the Treasury, plays a crucial role in safeguarding the financial system from illicit activity. Its primary mission is to combat money laundering, prevent the financing of terrorism, and promote national security through the strategic use of financial regulations and intelligence. FinCEN accomplishes this by collecting and maintaining vast amounts of financial transaction data, which it analyses and disseminates to support law enforcement efforts. As a key source of intelligence for tracking suspicious financial activities, including identity theft and check fraud, FinCEN provides critical insights into the scale and nature of financial crimes.
Classification Process
To understand the distribution of victimization across categories, we developed a Natural Language Processing (NLP) classifier, with the goal of categorizing the check issuer and check addressee based on the following categories: Business, Government, Nonprofit, Educational institution, Private, Religious, and Unknown.
To develop the classifier, we employed the NLP DistilBERT model. DistilBERT is a deep learning NLP model designed to consider context by analyzing the relationships between words in a sentence bidirectionally. To train the model we took a sub sample that included 6,608 checks with both issuer and addressee name. Those names were manually classified by a team of graduate students. To classify, each name was manually searched on Google, and if needed, we cross-checked with social media platforms such as Facebook, Instagram, and LinkedIn. The categorized data was inputted into the model, which converts the text into tokens that the model can understand.
The training process was carried out using the PyTorch library on python. The tokenized input from each fold was passed into a DistilBERT model fine-tuned for sequence classification. The model architecture was modified to output the correct number of classes, corresponding to the unique category.
To ensure robustness and avoid overfitting, the labeled data was divided into training and test sets, and a five-fold cross-validation approach was employed. The AdamW optimizer was used to adjust the model parameters, while the CrossEntropyLoss function calculated the loss between the predicted and actual labels. The model was trained for five epochs per fold, and during each epoch, the weights were updated based on the computed gradients.
After each fold, the model’s performance was evaluated on the validation set. Metrics such as accuracy, precision, recall, and F1 score were calculated to assess the model’s classification capabilities. The metrics from each fold were aggregated to provide an overall evaluation of the model’s performance across the dataset. The overall F1 score was 89%, indicating that the model was highly effective in correctly identifying and categorizing issuer and recipient entities, while minimizing both false positives and false negatives.
Once the cross-validation process was completed, the final model was trained on the entire set of labeled data. This final model was then used to make predictions on the rest of the dataset, where both business names and victim names was classified by the model. To further validate the model’s predictions, a random sample of 320 checks was manually labeled by a team of graduate students, demonstrating accuracy of 87%. This provided a reliable ground truth and allowed for fine-tuning and verification of the model’s predictive capabilities.
Results
RQ 1. How Does Check Theft Data Collected From Telegram Correspond With Official Reports of Check Theft Available Over FinCEN?
Over the course of the study, a total of 71,864 checks were analyzed. It is important to note that some checks were repeated across different Telegram channels, resulting in multiple appearances. After removing these duplicates, we obtained a total of 65,621 unique checks. Figure 2 illustrates the quantity of checks collected over each of these 16 months, both before and after removing duplicates. For each analyzed image, we extracted data pertaining to: the issuing bank’s name and address; the name and address of the individual or business to whom the check was issued; the name and address of the issuer; and the amount of the check.

Monthly count of checks analyzed via OCR, September 2021 to December 2022.
It is important to note that not all check images were clearly visible, and in some instances, attempts were made to obscure certain information deliberately. Consequently, some of the aforementioned data elements were unavailable for a subset of the checks.
Cross-referencing the distribution of stolen checks during the study period reveals notable differences between FinCEN reports and data from illicit markets. As shown in Figure 3, FinCEN reports indicate a sharp increase in stolen check incidents between November 2021 and April 2022, followed by a decline and a secondary peak in August 2022.

Monthly count of stolen check reported to FinCEN, September 2021 to December 2022.
In contrast, Figure 2 illustrates that the volume of stolen checks advertised on illicit markets follows a different monthly pattern, with peaks observed in October 2021, January 2022, and November 2022. These fluctuations do not align with the FinCEN timeline. Notably, the peaks in illicit market activity appear to precede those reported by FinCEN, suggesting a potential delay between when stolen checks are traded and when they are formally reported.
RQ 2. What Insights Can Check Theft Data Collected on Telegram Extend Beyond Official Data on Check Theft?
Amount Distribution
Examining the distribution of check amounts throughout the designated period. Out of the 65,621 unique checks analyzed, the amount was visible and documented for 64,629 of them. In addition to the OCR analysis, we manually reviewed and verified checks where the amount was either $10,000 or more, or $10 or less. The average amount across all checks was $4,944, while the median amount was $550. The discrepancy between the median and the mean is attributable to a few extreme cases involving high-value checks. Specifically, there were 594 checks with amounts exceeding $100,000, including 6 checks ranging from $900,000 to $1,000,000. Notably, more than 92% of the checks were below $10,000, and over 61% were below $1,000. Figure 4 presents the distribution of checks, with the final bin capturing amounts exceeding $12,000, while Figure 5 depicts the distribution of checks between $100,000 and $1,000,000. Additionally, Figure 6 shows the cumulative percentage distribution of checks by amount.

Distribution of check amounts.

Distribution of check amounts between $100,000 and $1,000,000.

Cumulative percentage distribution of check amounts.
Figure 7 shows the mean and median check amounts per month across the 16-month period. We can observe a positive trend in the average check amount over these months, with the average increasing fourfold, from approximately $2,500 in September 2021 to over $10,000 in September 2022, followed by a sharp decrease thereafter.

Mean and median check amounts per month.
While the mean and median provide a summary of central tendency, we do not report the mode due to the highly granular nature of the check amounts. Most values are unique or occur infrequently, resulting in no meaningful modal value. Nonetheless, the most frequently observed amount was $615, though it represents only a small fraction of the overall dataset.
States Distribution
Our analysis progressed by investigating the distribution of checks across all states throughout the study period. Of the 65,621 unique checks analyzed, state data was available for 31,167 checks. Table 1 presents an aggregated overview of the number of checks attributed to each state across all 16 months. The top states with the highest number of stolen checks identified were Florida, New York, Pennsylvania, Texas, and California, in this order. Each of these states had more than 2,400 checks, together representing approximately 60% of the entire check count. Notably, Florida, the state with the highest number, had a total of 5,352 checks.
Distribution of Checks by State.
Note. The numbers in parentheses indicate the percentage of each state’s contribution to the total number of checks.
Following this overview, Figure 8 displays a series of maps of the United States, each corresponding to a different month from September 2021 to December 2022, to illustrate the geographic distribution of checks. These visual representations make it evident that, across the majority of the months analyzed, the highest volumes of checks were consistently found in the aforementioned states.

Geographic distribution of checks by state, monthly overview from September 2021 to December 2022.
Figure 9 shows the number of checks for the top 5 states—Florida, New York, Pennsylvania, Texas, and California—in the left panel, and the mean and median check amounts in the right panel across the research period. As observed, there is no significant difference in the number of checks over time. However, there is an increase in the mean check amount from July to September 2022, while the median remained relatively stable throughout the period.

Number of checks and mean/median check amounts for the top 5 states over time.
We extend our analysis by examining the U.S. Department of the Treasury’s Financial Crimes Enforcement Network (FinCEN) report on check-related financial crimes from September 2021 to December 2022 and comparing the reported check fraud data with the information we collected from the Telegram groups. The FinCEN report provides the number of check fraud incidents by state during this period, with a total of 849,758 reported cases. Only states that had more than 1% of the checks in the Telegram groups are included in our analysis. Table 2 presents the percentage distribution of check fraud by state in the FinCEN database alongside our Telegram dataset, with a comparison between the two groups conducted using a series of two-proportion z-tests.
State-Level Comparison: FinCEN Versus Telegram Data.
p < .05 is indicated by **, and p < .01 is indicated by ***.
Among the top five states, we find that Florida, New York, Pennsylvania, and Texas are overrepresented in the Telegram data, while California is underrepresented. This may be due to the fact that California tends to report more cases of check fraud to authorities compared to the other states. Additionally, the composition of the Telegram group itself may be geographically biased, with more members concentrated in states like Florida, New York, Pennsylvania, and Texas, leading to a higher number of cases from those areas being shared, regardless of actual crime levels.
Banks Distribution
In our comprehensive dataset, we successfully obtained information on the originating bank for each of the 62,733 unique checks. We then focused on the nine largest banks in the United States by assets, calculating the percentage of total checks associated with these institutions. Overall, approximately 56.3% of the checks were linked to these banks. Figure 10 visually depicts the percentage of checks associated with each bank over the 16-month period covered by our study.

Percentage of checks associated with top U.S. banks, monthly analysis.
Additionally, we collected data for both states and bank names from 30,779 unique checks. In four out of the five top states previously mentioned (Florida, New York, Pennsylvania, Texas, and California), the percentage of checks associated with the top nine banks was higher than in the overall dataset. Table 3 shows the number of checks available for each of these states and the corresponding percentage linked to the top nine banks. Among these states, the top nine banks account for approximately 53% to 70% of the checks, with Florida having the highest concentration (70.0%) and Texas the lowest (52.7%). This suggests that in Texas, compared to other top states, a larger share of check theft activity appears to involve smaller financial institutions such as local or regional banks and credit unions.
Percentage of Checks from Top 9 Banks in the Top 5 State.
Lastly, Checks from the top nine banks were found to have significantly lower values compared to those from other banks. Specifically, our analysis revealed that the average value of checks issued by the top banks was approximately $4,363, whereas checks from other banks averaged about $5,645. Furthermore, the median check value for the top banks was around $495, in contrast to $615 for the other banks. A two-sample t-test and Wilcoxon rank sum test confirmed that these differences are statistically significant (p < .001).
Classification Distribution
As mentioned in the methodological section, the NLP classifier was developed to categorize check issuers and recipients into distinct categories, including Business, Government, Nonprofit, Educational Institution, Private, Religious Organization, and Unknown. The distinction between “N.A.” and “Unknown” in the table is important. N.A. (Not Available) indicates cases where the data was not accessible, either because it was deliberately obscured or the text was unreadable due to poor quality or missing information. On the other hand, Unknown refers to instances where the data was available, but the algorithm was unable to confidently assign a category, as the information did not meet the required threshold for accurate classification.
Out of the 65,621 checks analyzed, issuer data was visible in 32,466 checks, while recipient data was available in 52,643 cases. This discrepancy arises from the intentional obfuscation of certain check details in some Telegram channels and groups in order to conceal the details of the victims. Table 4 presents the number of checks in each category, separated for issuers and recipients, along with the average and median amounts recorded on the checks in each group.
Distribution of Check Counts and Monetary Values by Issuer and Recipient Categories.
We observe that among issuers, the largest group is the business sector, accounting for 56% of total checks. In contrast, among recipients, the two largest groups are the private sector (33.6%) followed by the business sector (6.6%). Since our research primarily focuses on stolen checks taken from recipients, we will concentrate our analysis on this group.
In the recipient group, categories such as Government, Nonprofit, Educational Institution, and Religious Organization are very small, each representing less than 0.4% of the total checks. This suggests that crime is predominantly concentrated on the business and private sector groups. Notably, the business sector exhibited substantially higher check amounts compared to the private sector. The average check amount in the business sector was $6,496 (median: $726), significantly exceeding the private sector’s average of $2,854 (median: $356). These differences were confirmed as statistically significant through both a two-sample t-test and Wilcoxon rank-sum test (p < .001).
We next examine whether the previously observed trends regarding check crime activity are also evident among our five top states, Florida, New York, Pennsylvania, Texas, and California—in terms of both the number of checks and their values among recipients. Table 5 presents the counts of checks for the business and private sectors in each state, along with the average and median amounts associated with these checks.
Distribution of Check Counts and Values by Sector in the Top 5 State.
Our analysis indicates that the percentage of checks classified under the business and private sectors closely aligns with the distribution observed in the overall check dataset. The private sector accounts for approximately 4% to 9.3%, while the business sector represents 36.5% to 46.9%. Furthermore, in every state analyzed, the check amounts in the business sector are significantly higher than those in the private sector.
Despite this consistency, significant disparities in check values across states are evident within both sectors. Notably, both business and private sector check values in Pennsylvania are considerably lower compared to the other four states. Furthermore, in New York and Texas, the average check value in the business sector is substantially higher, with mean (median) values of $15,677 ($1,001) in New York and $10,193 ($820) in Texas. In contrast, check values in the private sector remain relatively consistent across the top states.
Next, we examine the monthly percentage of checks in the private and business sectors over the 16-month period, along with their average and median values, as illustrated in Figure 11. Notably, the average and median values of checks in the business sector remain consistently higher throughout the entire period; however, we observe that the difference is more significant from Q1 to Q3 than in Q4, mainly because the average value of business checks is decreasing during this quarter. Simultaneously, the percentage of checks from the private sector reveals a significant downward trend, declining from 35.4% in September 2021 to just 22.3% by December 2022. While the percentage of business sector checks remains remarkably stable, the private sector exhibits a markedly different trajectory.

Monthly check percentage, averages, and medians trends for private and business sectors.
RQ 3. Does Check Theft Data Predict Identity Theft Data? Was It Always the Case?
Additionally, we test whether the Telegram dataset, which captures check fraud incidents as they are advertised on illicit markets, has a measurable effect on the FinCEN data, which represents officially reported cases of check fraud and identity theft. Given the inherent lag in the reporting process, victims often take time to realize their checks have been stolen and initiate formal complaints—we account for this delay by introducing a lagged analysis between the two datasets.
Our hypothesis is that trends observed in the Telegram data can serve as a leading indicator for trends in the FinCEN data, effectively capturing the early stages of check fraud activity before it is formally reported. To evaluate this, we examine whether variations in the volume of stolen checks advertised on Telegram are statistically significant predictors of subsequent variations in the FinCEN reports.
To test this hypothesis, we formulate fixed-effects models that account for variations across different months. These models include month-level fixed effects to eliminate unobserved variation constant within each time period, such as economic or seasonal factors. To achieve this, we present the following models:
where
A positive and statistically significant β coefficient in Model 1 indicates that an increase in the number of checks advertised on Telegram is associated with a corresponding rise in reported check fraud cases. In Model 2, a positive and significant β suggests that Telegram activity is similarly linked to an increase in reported identity theft incidents. In both models, β represents the expected change in FinCEN-reported cases for each additional stolen check observed on Telegram, controlling for time-specific effects.
The results for both models are presented in Table 6. For Model 1, which examines the relationship between FinCEN stolen check data and checks being sold on Telegram, we found a positive significant correlation across all lags. This indicates that higher appearances of stolen checks in Telegram groups corresponds to increase reporting of stolen checks in FinCEN data. The coefficients increase with longer lags, which may reflect that some victims take time to realize their checks have been stolen and report the theft to authorities. This highlights the value of Telegram data as a predictive indicator for future check fraud trends.
Results of Models (1) and (2): Telegram Data and FinCEN Reports.
Note. Standard deviations are shown in parentheses.
p < .001.
Model 2, which examines the relationship between identity theft cases reported in FinCEN and checks being sold on Telegram, shows a similar pattern. Positive significant correlations were found across all lags, with coefficients also increasing over time. The rise in identity theft reports may be partly driven by the stolen checks themselves, as cashing these checks often requires identity theft in some states and banks. This further supports our hypothesis, demonstrating a link between the sale of stolen checks on Telegram and subsequent increases in both check fraud and identity theft reports. However, the relatively low R² values in this model suggest that Telegram activity, while predictive, captures only part of the variation in reported identity theft.
RQ 4. What Kind of Check (Personal, Business, Government, or Other) is Most Likely to Result in Identity Theft?
To further explore the relationship between Telegram data and FinCEN reports we ran Model 1 and Model 2 to test whether trends observed in the Telegram data can serve as a leading indicator for trends in the FinCEN data, separately for checks issued to private and business sector victims. We hypothesized that both groups would show a positive relationship with the number of checks reported in the FinCEN data, but that only the private sector group would correlate with the FinCEN identity theft data. This hypothesis stems from the fact that identity theft is often required to cash out private sector checks, whereas business sector checks generally do not require such steps.
The results, reported in Table 7, supported our hypothesis. For Model 1, both private and business sector checks exhibited significant positive correlations with the FinCEN stolen check data across all lags, indicating that trends in Telegram activity predict trends in reported check fraud for both groups. However, in Model 2, only the private sector checks were significantly correlated with the FinCEN identity theft data. The lack of correlation for business sector checks aligns with the understanding that cashing these checks typically does not involve identity theft. These findings further highlight the distinct mechanisms of fraud between private and business sector victims and reinforce the predictive value of Telegram data for understanding trends in financial crimes.
Results of Models (1) and (2): Telegram Data and FinCEN Reports by Private and Business Sectors.
Note. Standard deviations are shown in parentheses.
p < .01. ***p < .001.
Discussion
This study provides a unique insights on the intersection between check theft and identity theft by leveraging OSINT data collected from Telegram illicit markets of stolen checks, an under explored yet increasingly central platform in the ecosystem of financial fraud. Our findings demonstrate that check theft activity observed on Telegram illicit markets is not only widespread and geographically concentrated but also predictive of officially reported check fraud and identity theft incidents. This underscores the potential of alternative data source to serve as early warning system for financial crimes and complement existing government reporting mechanisms.
A key finding of this analysis is the temporal relationship between check theft advertised on Telegram based illicit markets and subsequent identity theft reports field through FinCEN. The positive and significant correlations across 1, 2, and 3 month lags suggest that activity on Telegram illicit markets may act as a precursor to formal reporting of financial crimes. This has critical implications for the timeliness of fraud detection and response efforts. While official datasets offer verified, high-quality data, they are often constrained by delays, under-reporting, and limited granularity. In contrast, OSINT data, as can be seen in this analysis, can offer valid, real-time, high-volume insights that augment traditional data collection mechanisms.
Another important contribution is the differential impact of check type on the likelihood of associated identity theft. We find that checks issued to private individuals are more strongly associated with subsequent identity theft reports than those issued to businesses. This distinction is crucial for tailoring risk assessments and fraud prevention strategies. Since private checks often necessitate identity impersonation for cashing or laundering, they present a higher threat vector and thus require more robust protective measures.
While this study offers significant insights, several limitations should be acknowledged. First, data collection was limited to Telegram and certain online illicit markets, which may not represent all channels used for check fraud. Additionally, the use of OCR technology, while instrumental in data extraction, is limited by the quality of images and may have introduced errors. We recommend that future studies address these limitations by expanding data sources and incorporating additional validation steps, such as manual checks on a broader subset of OCR-processed data to improve accuracy. Furthermore, the 16-month timeframe may not capture long-term or seasonal trends in check fraud activity, and future research could extend the observation period to determine if patterns fluctuate based on economic conditions or other factors.
Nevertheless, these findings yield several policy implications. First, the findings suggest that integrating OSINT into law enforcement intelligence frameworks. Given the predictive power of Telegram-based check fraud data, federal and state-level financial crime units should consider incorporating OSINT monitoring into their surveillance protocols. Automated scraping and classification of content from encrypted or semi-open platforms, combined with machine learning techniques like those applied in this study, can generate timely alerts on emerging fraud patterns. Developing centralized OSINT monitoring hubs under agencies such as FinCEN or the USPS Inspection Service could improve response times and resource allocation.
Second, enhancing private public-private data sharing mechanisms. Banks, check-processing firms, and digital security companies possess valuable proprietary data on suspicious transactions and account behavior. Facilitating structured data-sharing partnerships between these institutions and law enforcement, especially when OSINT indicators signal increased risk, could close information gaps and foster proactive mitigation. Encouraging secure, anonymized data exchange can also support more robust cross-validation of OSINT-derived signals.
Third, development of targeted prevention and victim support programs. The elevated risk of identity theft among private check recipients calls for tailored interventions. Government agencies and nonprofit organizations should focus on at-risk populations—especially older adults and low-income individuals, by offering fraud education campaigns, identity protection services, and support for recovering stolen funds or repairing credit histories. Victim assistance programs should be scaled to reflect the growing volume of check- and mail-related fraud.
While our findings underscore the utility of OSINT in identifying trends and predicting identity theft, particularly through analyzing publicly available data on encrypted platforms like Telegram, institutionalizing such surveillance methods demands a rigorous examination of ethical and privacy implications. To balance proactive fraud detection with the preservation of privacy, we argue for a multi-layered set of safeguards. First, transparent governance is essential: institutions should establish clear policies defining acceptable OSINT uses, data retention limits, and public disclosures of methodologies (where security considerations allow), as recommended by the Association of Internet Researchers (Markham & Buchanan, 2012). Second, privacy-preserving technologies, such as anonymization protocols and differential privacy, can reduce the risk of re-identifying individuals (Kosinski et al., 2013).
Third, proportionality and oversight must guide OSINT operations: independent oversight boards or ethics committees should ensure that data collection is justified, proportionate, and rights-respecting. Fourth, regulatory alignment is needed, since existing laws like the GDPR remain ill-equipped to address the nuances of public data when repurposed for surveillance (Wachter et al., 2017). Finally, user empowerment through digital literacy programs can help individuals better control their online footprints, reducing unintended exposure. By embedding these safeguards and fostering collaboration among policymakers, technologists, and civil society, institutions can harness OSINT’s predictive power without sacrificing user privacy.
Conclusion
This study demonstrates the feasibility and utility of integrating OSINT into the broader toolkit for combating financial fraud. By examining over 65,000 unique stolen checks circulated on Telegram and mapping them against official reports, we show that online illicit market activity can reveal early signals of identity theft and fraudulent check usage. While the challenges posed by encrypted platforms are significant, they are not insurmountable. With the right regulatory frameworks, technological investments, and collaborative networks, stakeholders can shift from reactive to proactive financial fraud prevention. In doing so, we can better protect individuals and institutions from the growing threat posed by check theft and identity misuse in the digital age.
Footnotes
Funding
The authors received no financial support for the research, authorship, and/or publication of this article.
Declaration of Conflicting Interests
The authors declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
