Groundeddocs
ReferenceWeb sources and crawling

Fetch one page of a web source again now (editors and above)

POST/v1/teams/{team}/sources/{sourceId}/documents/{documentId}/refetch

Starts a crawl run (trigger "page") of just this page's URL: it is fetched, stored and indexed again if it changed. Links aren't followed and no other page is touched or removed, and the source's schedule doesn't change. The run counts toward the team's crawl limits like any other. Only pages of web sources can be re-fetched (400 not_web_document); a source with an active run answers 409 crawl_in_progress.

Path Parameters

team*string

Team slug or ID

sourceId*string
Formatuuid
documentId*string
Formatuuid

Response Body

application/json

application/json

application/json

application/json

application/json

application/json

curl -X POST "https://example.com/v1/teams/string/sources/497f6eca-6276-4993-bfeb-53cbbbba6f08/documents/497f6eca-6276-4993-bfeb-53cbbbba6f08/refetch"
{  "data": {    "id": "497f6eca-6276-4993-bfeb-53cbbbba6f08",    "sourceId": "797f5a94-3689-4ac8-82fd-d749511ea2b2",    "status": "queued",    "trigger": "create",    "pagesDiscovered": 0,    "pagesFetched": 0,    "pagesChanged": 0,    "pagesUnchanged": 0,    "pagesSkipped": 0,    "pagesFailed": 0,    "documentsDeleted": 0,    "truncated": true,    "truncatedReason": "max_pages",    "waitingReason": "concurrent_crawls",    "waitingUntil": "2019-08-24T14:15:22Z",    "error": "string",    "createdAt": "2019-08-24T14:15:22Z",    "startedAt": "2019-08-24T14:15:22Z",    "finishedAt": "2019-08-24T14:15:22Z"  }}