The cloud broke DLP's old model. When data lived in a handful of on-prem systems, you could integrate with each. When it lives in hundreds of SaaS apps — and a new one onboards every week — building a bespoke DLP connector per app is a treadmill no vendor can win.
Netskope, Inc.'s US11575735B2, “Cloud application-agnostic data loss prevention (DLP)” (issued February 7, 2023; CPC H04L 67/10 — cloud/distributed computing, and G06F 9/547 — remote procedure call/RPC), describes DLP that is agnostic to the specific cloud application. Read it at US11575735B2.
“The technology disclosed applies data loss prevention (DLP) to those cloud-applications for which no application-specific parser is available.”— U.S. Patent No. 11,575,735 source
The architecture starts with an inline proxy. The specification places “a proxy positioned between the endpoints and the cloud-based services” that “intercepts and parses a message between an endpoint and a server,” determines “which cloud-based service application programming interface (API) is being accessed,” and applies a parser — what it also calls a connector — “to collect metadata.” The whole problem reduces to: when there is no parser written for this specific app, what do you do? The filing's answer is a fallback hierarchy of parsers rather than a per-app build.
The clever middle tier is the category-directed parser. The filing groups known services into categories — it names “personal pages and blog,” “news websites,” “cloud-based storage services,” and “social media services” — each defined as a list of provider URLs whose users “perform similar activities.” Providers in a category use different syntaxes, so the category-directed parser “synthesize[s] interaction syntax patterns of a sample of providers in the category” and collects metadata using “multiple category-directed match rules synthesized from syntaxes used by the sample providers.” The insight: apps in the same category move data in structurally similar ways, so one set of rules learned from a few representatives generalizes to the rest — including apps the vendor never specifically integrated.
When even a category does not fit, a third tier takes over. Claim 1 walks the decision: the system determines a cloud service “is being accessed via an application programming interface,” determines “that no service-specific parser is available,” determines “a category of service,” and selects “either a category-directed parser based upon the determined category or, if no category-directed parser is available… a generic parser” built from “at least one default match rule.” Metadata that any tier collects “enables the DLP processor to focus analysis of the content being conveyed via a corresponding API.” One policy engine thus covers the long tail of SaaS by reasoning about cloud data movement generically — service-specific parser if it exists, category parser if not, generic parser as the floor.
The dependent claims put numbers on what “synthesized from a sample of providers” actually means — and the numbers are the moat. A category-directed parser's match rules are not a token sample: claim 3 requires “at least ten match rules,” and claim 4 requires they be “derived from syntaxes used by at least twenty five providers” of the same category. Claim 5 requires the system to field “at least five category-directed parsers in distinct categories,” and claim 6 names them: “personal pages and blogs, news websites, cloud-based storage services, webmail, and social media services.” The routing decision itself is lightweight — claim 2 selects the category-directed parser “based on a domain name in a” URL. So coverage of an app the vendor never integrated comes from the domain pointing it into a category whose rules were generalized across dozens of similar providers. That is the engineering behind the breadth claim a buyer cares about.
The system's reach also depends on getting traffic to the proxy in the first place, and the specification handles both ends of that. In a “managed device” implementation, endpoints “are configured with routing agents… which ensure that requests for the cloud-based services… are routed through the inline proxy… for policy enforcement.” But it also covers the harder case: in an “unmanaged device” implementation, endpoints with no routing agent “can still be under the purview of the inline proxy… when they are operating in an on premise network monitored by the inline proxy.” Metadata the parsers collect lands in a “metadata accumulation store,” and category membership is looked up in a “category database.” The architecture is thus complete in two dimensions at once — it captures traffic from both managed and unmanaged endpoints, and it parses that traffic whether or not a purpose-built connector exists. Coverage breadth is the product of both: every device, every app, one policy engine.
Why this is a business story: app-agnostic coverage is the scalability argument at the heart of the cloud-access-security-broker (CASB) and SASE pitch, and it is central to Netskope's positioning as it moved toward the public markets. The competitive question every buyer asks — how many of my SaaS apps does this actually cover? — is answered precisely by the category-and-generic fallback this filing describes: coverage no longer depends on the vendor having pre-built a connector for your exact app. Netskope, Zscaler, and Palo Alto all filed heavily in cloud DLP in this window, signaling a high-stakes feature race over exactly this breadth claim.
The grounded read: app-agnostic DLP runs every cloud session through an inline proxy and falls back from service-specific to category-directed to generic parsers — learning shared syntaxes so one engine covers SaaS apps it was never built for. Netskope's 2023 grant names that tiered-parser approach in claim-level detail — the scalability claim underpinning the modern CASB and SASE pitch.
Comments
Loading comments…