By Varunavi Bangia
In the digital age, technology has ushered in an era of unprecedented convenience and innovation, yet it has simultaneously created a landscape of new, systemic vulnerabilities. As regulatory frameworks struggle to keep pace with rapid technological advancement, a significant disconnect has emerged between legal safeguards and the reality of how the modern data economy functions. At the heart of this tension lies a foundational misunderstanding: our current laws are built on the protection of the individual, while the data economy has shifted its focus toward the exploitation of groups.
The Myth of Individual Consent in a Big Data World
The prevailing paradigm for data protection across most global jurisdictions is rooted in the concept of "informed consent." Regulators generally assume that if an individual agrees to data collection at the point of origin, their privacy is adequately managed. However, the data economy has evolved far beyond the simple monetization of personal data as collected.
Today, the engines of Big Data, inferential analytics, and knowledge discovery operate on a different plane. These technologies do not merely process the data an individual provides; they generate entirely new insights—inferences about health, behavior, political inclination, and financial status—that the individual never explicitly consented to share. Scholars have identified this shift as the "de-individualization of the person." Instead of treating a user as a distinct human being with unique traits, platforms now categorize individuals as mere members of a specific data cluster.
This profiling is increasingly invasive. By aggregating massive, diverse datasets, corporations and governments can build digital dossiers that are far more revealing than any single stream of information. Furthermore, even when data is technically anonymized, it can often be cross-referenced with secondary datasets to enable re-identification. In this environment, the most significant harm to privacy is no longer just the exposure of private data, but the discriminatory use of inferential data to categorize, stigmatize, or exclude individuals from essential opportunities.
The Regulatory Gap: Why Current Definitions Fail
The major challenge with current data protection frameworks—such as the Data Protection Bill (2021) and its predecessors—is the restrictive definition of "personal data." Current law typically defines personal data as information relating to a natural person who is directly or indirectly identifiable. This narrow focus effectively excludes anonymized data from the domain of regulation.
Once data is anonymized or generalized, it falls outside the scope of protection. Yet, the impact of such "non-personal" data usage on an individual remains profoundly critical. If a corporation uses an anonymized dataset to identify a community as "high-risk" for a particular disease, it may deny insurance or employment to every individual within that demographic, often without the person ever knowing the reason for their rejection.
Moreover, while some inferences are legally considered "personal data," the scope of inferential profiling is much broader. Profiling can occur without ever using personally identifiable information. Because the law limits the definition of profiling to the processing of personal data, an entire class of predictive analytics remains unregulated. To address this, legal experts argue that inferential data should be categorized based on whether the inference itself can identify a person or a group, or whether it relates to a sensitive attribute, rather than focusing solely on the raw data used to generate the inference.
Chronology of Regulatory Inadequacy
- The Rise of Big Data (Early 2000s): The emergence of Knowledge Discovery in Databases (KDD) challenged the traditional legal focus on individualistic privacy.
- The Puttaswamy Verdict (2017): In the landmark Justice K.S. Puttaswamy v. Union of India case, the Indian Supreme Court explicitly recognized that informational privacy is a fundamental right, acknowledging that aggregation of seemingly non-private data can create detailed, invasive profiles.
- The Srikrishna Committee (2018): Despite the Court’s warnings, the resulting reports and subsequent iterations of the Data Protection Bill failed to address the specific dangers posed by big data analytics and group-based profiling.
- The Data Protection Bill (2021): This iteration maintained a narrow definition of personal data, effectively ignoring the economic and social value that corporations derive from non-personal data when creating group profiles.
Categorizing Data for the Future
The binary distinction between "personal" and "non-personal" data is increasingly untenable. Regulators are currently grappling with how to treat non-personal data, often viewing it purely as an untapped economic resource to be harnessed for development. This creates a false dichotomy: the belief that privacy protection must inherently come at the cost of economic growth.

To bridge this gap, we must move beyond the mutually exclusive treatment of these data types. A more effective regulatory approach would be to categorize non-personal data into two distinct heads:
- Human Non-Personal Data: This includes anonymized personal data that can still be used, directly or indirectly, to identify a person or a group. This should be re-classified and regulated under the same stringent standards as personal data.
- Non-Human Non-Personal Data: This encompasses data that is non-human in nature at the time of collection and in its use (e.g., weather patterns, sensor data from non-living objects).
Moving Beyond the Individual: A New Concept of Privacy
The most radical, yet necessary, change required in our legal framework is a shift from protecting the individual to protecting the group. In the digital age, group privacy is not merely an extension of individual privacy; it is a distinct, collective right.
Current stewardship models—which rely on group loyalty or organizational structures—are insufficient because the groups being profiled are often "manufactured" by algorithms. These are clusters of people who have no prior connection to one another, brought together solely by shared, inferred behaviors.
The violation of group privacy occurs when an individual suffers a consequence—such as being denied a loan or facing systemic bias—solely because of their inclusion in an algorithmically generated group. Therefore, the remedy must be collective. We need to move toward a concept of "Categorical Privacy," which applies to aggregated data that may no longer identify an individual but still carries the potential for discriminatory impact.
Implications and The Path Forward
The path to a more secure digital society requires three fundamental regulatory interventions:
- Ex-Ante Impact Assessments: Regulators must mandate that entities conduct "privacy and social impact assessments" before deploying profiling technologies. These assessments should audit algorithms for potential bias and discrimination, moving beyond mere technical compliance.
- Hard Accountability and Auditing: There must be a move toward robust oversight where algorithmic decisions can be audited. If an algorithm systematically discriminates against a specific group, the organization must be held accountable for the social consequences, regardless of whether they intended to target individuals.
- Jurisprudential Realignment: The theoretical basis for group privacy already exists within the Indian constitutional framework, as evidenced by the Puttaswamy and Navtej Johar cases. The judiciary has already recognized that the aggregation of data is a threat to constitutional values. It is now up to the legislature to translate these judicial warnings into concrete, enforceable laws.
Conclusion: A Call for Unified Protection
The ongoing effort to treat personal and non-personal data as separate domains is a failure of imagination. In the age of big data, the line between an individual’s identity and their group association is blurred by design. As long as our laws focus solely on the "identifiable individual," they will remain powerless to stop the most insidious harms of the digital age: the automated, systematic, and often invisible discrimination of groups.
True data protection requires a sophisticated, holistic approach—one that understands that in a world driven by predictive analytics, protecting the group is the only way to effectively protect the individual. The coming decade will prove that our ability to regulate these technologies will define not just our economic future, but the health and equity of our society at large.
Disclaimer: This research was supported by an unrestricted scholarship pursuant to the Facebook India Tech Scholars Program 2021-2022. The work was conducted under the research mentorship of the Software Freedom Law Center and without any oversight from Meta. The views expressed herein are those of the author and are not necessarily those of Meta or the Software Freedom Law Center.
Varunavi Bangia is a 5th-year student of BA LLB (Hons) at the West Bengal National University of Juridical Sciences (WBNUJS), Kolkata.

