Govern Data Collections and Sources
Administrators use Data to govern the boundary between external systems, stored credentials, indexed content, collection access, and the agents that use that content. A source working technically is not enough: it must also have a clear owner, approved scope, refresh policy, and retirement path.
The Data Guide
- Configure data upload limits to establish sensible defaults and per-user exceptions before ingestion begins.
- Configure enabled connections to control which integrations are approved for your organization.
- Choose the best data source for your use case before storing or synchronizing business content.
- Design a data collection architecture for your teams so ownership and access match how your organization operates.
- Create and populate your first data collection using the approved scope and access model.
- Monitor data-storage utilization to track consumption and address capacity risks early.
Use these pages as the implementation checklist; use this Admin Guide to define identity ownership, least privilege, collection access, review, monitoring, and retirement.
For every source, keep these objects separate:
- Connection: authentication and provider identity.
- Data source: the folder, path, blob selector, share, project, space, or URL being ingested.
- Collection: the access-controlled business grouping.
- Agent assignment: which production behaviors can retrieve from the collection.
Admin Design Checklist
Before creating a production source, decide:
- Who owns the upstream content?
- Which service or user identity will authenticate?
- Which exact folders/projects/spaces/paths are in scope?
- Who may use the resulting collection?
- Who may edit, move, synchronize, or delete its sources?
- How quickly must upstream changes appear?
- How will you test retrieval quality and missing-answer behavior?
- What happens when the credential owner leaves or the source is retired?