This describes what CiteraHub collects and why. It is a draft and has not been reviewed by a lawyer.
What is collected about you: your email address, which is how you sign in; the organizations you belong to and your role in each; and a record of what you approved and when, which exists because an approval that cannot be evidenced is not an approval.
What is collected about your sites: the pages CiteraHub fetched, the response headers it saw, and the files it read from a repository you connected. This is what a finding cites, and a finding without its evidence is an assertion.
What is deliberately not collected: the content of a page is not sent to a model provider, and the default for that setting is off. Telemetry is off by default. Neither is turned on by a deployment choosing not to think about it.
Secrets are never stored in a report or a log. Where a credential-shaped value is found in a repository, its name and location are reported and its value is not read into anything.
Connector tokens are encrypted with a key bound to the tenant, the account and the provider, so ciphertext moved between tenants fails to decrypt rather than yielding the wrong token.
Retention defaults to ninety days for audit data. Personal data is deleted on request, and that deletion reaches the record of what you agreed to, because an erasure obligation is not answered by an audit trail.
Sub-processors, international transfers, the legal basis for each purpose, and the rights available to you under any particular jurisdiction are not drafted here. A qualified reviewer must write them.