veille Product, Discovery
AI System Discovery

You can't comply with systems you don't know you're running.

Veille reads six independent sources to find the AI in your organization: what you pay for, what your people reach, what they authorized, what is installed on their machines, what your automations call, and what somebody issued a key for. Each one catches something the others structurally cannot see.

Detection channelsSpend, network, OAuth, devices, automations, keys
Key obligationLaw 25 Art. 3.1, owner per system
Scan modeRead-only, no write access

The problem

Shadow AI is the most common source of Law 25 exposure.

Under Law 25, every AI system processing personal data in Québec must have a designated responsible person (Art. 3.1) and, for systems making automated decisions, a transparency notice (Art. 12.1) and a Privacy Impact Assessment (EFVP, art. 3.3 and 17).

The systems most likely to lack these are the ones no one thought to register: the HR chatbot an operations team bought on a credit card, the pricing experiment an ML engineer deployed for a sprint and never decommissioned, the meeting note-taker an employee connected to the company calendar with a personal account.

Shadow AI is not primarily a security problem. It is a documentation problem. An undeclared system cannot have a designated owner, cannot have an assessment on file, and cannot appear in a model inventory. Those absences are exactly what a regulator sanctions.

What Veille measures

Veille does not publish an industry average for how much shadow AI an organization has, because we have not measured one. What the product does is compute your figure: it reconciles the systems your teams declared against the systems the three channels actually found, and reports the difference as a coverage rate you can act on and re-measure. That number is derived from your data at scan time, and it is frozen in the audit trail so it can be shown later.

6
Independent detection channels. A system found in more than one is corroborated, not counted twice, and every corroborating channel raises its confidence score.
Art. 3.1
Law 25 requires a designated person in charge (Art. 3.1); every AI system still needs an identifiable owner accountable to them
$10M
Maximum administrative monetary penalty (or 2% of worldwide turnover) for failure to comply with privacy officer and documentation obligations. The penal fine reaches $25M or 4%.

Coverage

Six sources, six different blind spots.

Most AI inventories are built from procurement data alone. That is why they miss the usage that carries the most risk: the tool nobody bought.

Channel 01
Software spend

An export from your accounting or expense system. Finds the AI your teams already pay for, including AI bundled into licences you already hold: Microsoft 365 Copilot, Salesforce Einstein, Notion AI, Zoom AI Companion. Matched on the AI-specific product name, so a plain Microsoft 365 or Sales Cloud line does not trigger a detection.

Blind spot: anything free.

Channel 02
Network logs

A domain export from your firewall, proxy or DNS resolver. This is the only channel that sees free-tier AI used on personal accounts, which has no invoice, no licence and no admin console entry. Blocked requests are reported too: an attempt is still evidence of intent.

Privacy: results are aggregated per vendor. User identifiers are counted, never retained. Per-user attribution stays off unless you deliberately enable it.

Channel 03
OAuth grants

The third-party applications your people authorized with "Sign in with Microsoft" or "Sign in with Google". Uniquely, this channel also reports what data each one can read. A note-taker holding calendar access is an integration; the same note-taker holding mailbox and full document access is a different finding entirely, and only one of them belongs at the top of a remediation queue.

No credentials required: both admin consoles export this list today.

Channel 04
Devices and browsers

An export from your device management, or a browser-extension report. This is the only channel that sees software running locally: a desktop transcription app bills nobody, appears in no admin console, and a model running on the machine makes no outbound request for a network log to catch. It also finds the highest-exposure case there is, an extension holding "read and change all your data on all websites", which sits inside every authenticated session that employee has.

Privacy: only AI software is retained; everything else is dropped, so this cannot become a general software census. Device identifiers are counted, never retained, unless you deliberately enable it.

Channel 05
Automations

A workflow list from Zapier, Make, n8n or Power Automate. Every other channel finds a tool a person uses; this one finds a tool a process uses. A workflow that summarises inbound CVs with a model processes personal information at 3am on a Sunday with no employee present, and on Microsoft's platform it arrives inside a licence you already hold, so your spend export shows only a line that says Microsoft 365.

Metadata only: the workflow's name and its connector list. Never its field mappings, its data or its prompt text.

Channel 06
Keys and credentials

A name list from your secret manager or your CI variables. A credential is intent made concrete: somebody created an account and issued a key so software could call a model unattended. That arrives earlier than an invoice, because trials and free tiers bill nothing. It also finds the key issued once, pasted into a service and forgotten, where there is no recurring charge and the person who created it may have left.

No secret value is ever read. Not to fingerprint it, not to check what it looks like, not to test whether it still works. Export the names; if you send a value anyway, it is stripped before anything reads the row.

How it works

The discovery process.

01
Guided intake, declare what you know
The onboarding starts with a structured intake form. Your team declares the AI systems they are aware of, their purpose, the data they process, and who owns them. This establishes the declared baseline that everything afterwards is measured against.
02
Multi-channel scan, find what you don't know
Veille reads the sources above and resolves each finding against a versioned catalog of known AI vendors, covering foundation-model APIs, cloud ML platforms, AI embedded in ordinary SaaS, meeting intelligence, developer assistants, and the higher-risk categories: HR screening, credit decisioning, biometric identity, health. Every scan is read-only, with no write access and no data extraction. Detection is pattern-based and deterministic, with no LLM call, so a scan cannot hallucinate a system that does not exist. Connectors for MLflow, Databricks, GitHub and SageMaker are in beta, read-only, metadata only.
03
Reconciliation, quantify the gap
Veille compares what was declared against what was found and reports four figures: declared, detected, never declared and without an owner, and the share of your inventory that is actually governed. A vendor found in several channels is recorded as one system with corroborating evidence, not as duplicates. A scan that finds nothing is recorded too, because "we scanned and found nothing" is a defence and "we never scanned" is not.
04
Governance, close each gap
Every discovered system enters the registry as ungoverned, with a suggested risk tier and the frameworks likely to apply. Nothing is auto-concluded: detection surfaces the finding, a human confirms the responsible person. Once an owner is confirmed, the system becomes governed and is scanned against its applicable obligations. Re-running discovery on a cadence surfaces what appeared since the last pass, and decommissioned systems are logged as a compliance event in the audit trail.
05
A finding has a life, not a label
A detection is not a verdict. Each finding moves through nine recorded states — detected, under investigation, confirmed, registered, approved, restricted, prohibited, false positive, closed — and every move is written to the tamper-evident audit trail with who made it and when. That is what lets you answer the questions an auditor actually asks: when did you first see it, who looked at it, was it still in use after you told them to stop. A finding dismissed as a false positive has to carry a reason and the name of the person who wrote it, and it can be reopened when new evidence arrives.
06
Confidence you can argue with
Every finding shows the arithmetic behind its score, channel by channel: how far each source is trusted on its own, how many times it saw the tool, how much its age cost it. Independent channels accumulate; two observations from the same log do not, because they are one witness. Below the confirmation threshold the product says so in words — a lead, not a finding — and refuses to let it be confirmed. A reviewer can also set aside one signal as noise without discarding the others, which is what stops a single wrong hit from burying three correct ones. Nothing here ever reads 100 %: it is all inferred from observation, and a displayed certainty invites people to stop asking.
Scheduling, stated plainly

Scans are operator-triggered today, from an upload or a stored read-only connector credential. Scheduled recurring scans are built and tested but are not yet enabled in a hosted deployment. We would rather tell you that than describe a cadence you would not actually get.

Question

How many AI systems does your organization actually have?

Book a 30-minute call. Bring one export, from any of the three channels, and we will run it live and show you the reconciliation on your own data.

Book a call