
Large-scale unauthorized scraping has become one of the defining security challenges of AI. As automation grows more sophisticated and the demand for training data intensifies, a new class of security work has emerged: not just defending systems from intrusion, but protecting the data that lives inside them from being quietly drained at scale.
For Malya Jain, a security engineer who has spent over a decade in cybersecurity, this challenge sits at the center of her mind. With 13 years of continuous work across consulting and security incident response, she has now moved to what she describes as a genuinely new discipline: anti-scraping.
An emerging risk to data
More and more digital platforms are using AI to improve their internal operational efficiency. However, the models behind those systems need substantial, high-quality training data to deliver trustworthy outputs. This growing demand for training data has elevated web scraping to a major challenge in the security and privacy spaces.
In technical terms, unauthorized scraping occurs when systems automatically and systematically extract information from sources like web pages, APIs, or public interfaces. The current methods used for this have greatly advanced and can look a lot like everyday browsing behavior, meaning there’s a challenge to tell apart real user activity and automated access—let alone determining whether such automation is authorized or malicious.
As the years have passed, these techniques have only become more accessible and easier to deploy, putting methods within reach of people with little to no technical knowledge. In fact, research shows automated traffic makes up a substantial portion of web activity, with over half of all traffic originating from bots.
The implications of this go further than the usual security worries. Unauthorized large-scale data collection can present serious risks to data privacy, and it can also compromise a company’s ability to meet strict security standards for its customers as well as for meeting industry regulations.
For Malya Jain, this growing trend is only going to continue expanding, and as such, should be treated as a core issue within cybersecurity. “Platforms with user content are data-rich environments,” she notes, “and protecting that data has become integral to core security operations.”
Jain’s cybersecurity expertise
Jain’s path through cybersecurity has tracked the field’s own evolution almost in parallel. Early in her career, she worked across IT application controls and general security assessments, gaining broad exposure to how different organizations handle data integrity.
Later moving into security incident response, triaging cybersecurity events and resolving them. Each chapter reinforced the same underlying reality: the threat surface was always growing, and the people trying to defend it had to grow faster.
“In the past 13 years, the landscape of cybersecurity has expanded exponentially,” she says. “And it’s not slowing down. If anything, the integration of AI is accelerating everything.”
Privacy as a non-negotiable
One of the themes Jain returns to frequently is how much public and regulatory attitudes toward data have changed, and how consequential that shift has been for the security field.
“Five or six years ago, people didn’t really care if their email account had all their data, or if they were being tracked,” she says. “But now people are genuinely concerned. The awareness around data privacy has changed completely.”
Regulations like GDPR accelerated that shift, establishing a new standard for how organizations must handle user data and creating accountability structures (including audits and significant financial penalties) that didn’t previously exist at scale. For security teams, this changed the stakes of the work. Privacy and compliance are no longer separate concerns from security; they’re inseparable from it.
Jain sees this as a net positive, even if it raises the bar considerably. Greater awareness means greater accountability, and accountability, she argues, is what drives the kind of serious, sustained investment in security infrastructure that users deserve.
Leveraging AI on both sides
Jain outlined what she sees as the central tension now defining the field: AI is simultaneously empowering attackers and enabling defenders, and the race between the two is tightening.
On the attacker side, AI has dramatically compressed the time and skill required to execute sophisticated threat actor operations. Tasks that once took hours of manual scripting and experimentation can now be completed in a fraction of the time, with AI tools guiding users toward novel attack vulnerabilities that might not have been obvious otherwise. “Threat actors are finding more novel ways to perform digital attacks with AI’s help,” Jain says plainly. “It’s helping them do more, faster.”
What makes this particularly challenging is that modern attacks are increasingly designed to blend in. Instead of hitting a system with high-volume, obviously automated requests, sophisticated actors now employ strategies that deliberately pace activity to mimic human behavior patterns, mimicking the subtle irregularities of normal usage to avoid triggering detection systems.
“Attack systems can now imitate ordinary user patterns with enough accuracy to resemble legitimate activity,” she explains. “That means detection systems have to be sharp enough to spot behavioral anomalies that would slip past manual analysis entirely.”
On the defensive side, the same AI capabilities are being applied to build smarter, more adaptive detection systems that learn from behavioral patterns rather than relying on fixed rules or predefined signatures. Jain characterizes this as an evolving feedback loop: sophisticated automated attacks prompt advanced AI-driven defenses, which in turn shape the development of next-generation security capabilities.
The case for industry-wide collaboration
One conclusion Jain has reached through her years in this space is that no organization can solve this problem alone. The scale and sophistication of digital attacks require a level of coordination across the industry.
She advocates for shared security frameworks, more open channels between platform security teams, academic researchers, and penetration tester communities, and greater consistency around responsible data practices. The intelligence required to stay ahead of emerging threats, she argues, is too fragmented across silos to be effective when kept isolated.
“Use of AI to create automated scripts and attacks presents challenges that are felt across the entire digital ecosystem,” she says. “The response has to match that scope.”
The importance of guardrails against data harvesting
The rise of AI-driven and guided digital attacks has forced technology companies to evolve their entire conception of what a threat looks like. For security leaders trying to navigate this new paradigm, Malya’s perspective is a useful reminder that the next main challenge in the world of cybersecurity will consist of finding the balance between the growing use of automation and the need for data integrity.
Disclaimer: GeekWire newsroom and editorial staff were not involved in the creation of this content..