Skip to main content

Govern Data Collections and Sources

Administrators use Data to govern the boundary between external systems, stored credentials, indexed content, collection access, and the agents that use that content. A source working technically is not enough: it must also have a clear owner, approved scope, refresh policy, and retirement path.

The Data Guide

  1. Configure data upload limits to establish sensible defaults and per-user exceptions before ingestion begins.
  2. Configure enabled connections to control which integrations are approved for your organization.
  3. Choose the best data source for your use case before storing or synchronizing business content.
  4. Design a data collection architecture for your teams so ownership and access match how your organization operates.
  5. Create and populate your first data collection using the approved scope and access model.
  6. Monitor data-storage utilization to track consumption and address capacity risks early.

Use these pages as the implementation checklist; use this Admin Guide to define identity ownership, least privilege, collection access, review, monitoring, and retirement.

For every source, keep these objects separate:

  • Connection: authentication and provider identity.
  • Data source: the folder, path, blob selector, share, project, space, or URL being ingested.
  • Collection: the access-controlled business grouping.
  • Agent assignment: which production behaviors can retrieve from the collection.

Admin Design Checklist

Before creating a production source, decide:

  • Who owns the upstream content?
  • Which service or user identity will authenticate?
  • Which exact folders/projects/spaces/paths are in scope?
  • Who may use the resulting collection?
  • Who may edit, move, synchronize, or delete its sources?
  • How quickly must upstream changes appear?
  • How will you test retrieval quality and missing-answer behavior?
  • What happens when the credential owner leaves or the source is retired?