# Bring Back and Improve Scheduled Scraping for Knowledge Files (with Webhook Support)

**URL:** <https://community.pickaxe.co/t/bring-back-and-improve-scheduled-scraping-for-knowledge-files-with-webhook-support/6733>\
**Category:** Feature Requests\
**Created:** [July 31, 2025, 4:35am UTC](https://community.pickaxe.co/t/bring-back-and-improve-scheduled-scraping-for-knowledge-files-with-webhook-support/6733 "2025-07-31T04:35:45Z")\
**Posts on this page:** 6\
**Page:** 1

<div class="post-metadata">

**Author:** ![taedog2020](https://yyz2.discourse-cdn.com/flex004/user_avatar/community.pickaxe.co/taedog2020/32/809_2.png) [@taedog2020](https://community.pickaxe.co/u/taedog2020)\
**Post date:** [July 31, 2025, 4:35am UTC](https://community.pickaxe.co/t/bring-back-and-improve-scheduled-scraping-for-knowledge-files-with-webhook-support/6733/1 "2025-07-31T04:35:45Z")

</div>

**Summary:**  
Reintroduce and upgrade the ability to **schedule website scraping** for Knowledge Files. Enhance it with better control, visibility, and webhook support—making it a reliable, flexible system for keeping AI agents synced with live external data.

* * *

**The Problem:**  
Today, scraping a URL into a Knowledge File is a one-time event. If the source changes, the data goes stale and must be manually refreshed. That’s unsustainable for creators using Pickaxe with:

- Dynamic web content (blogs, SOPs, live docs)

- AI agents that depend on accuracy

- Workflows that rely on data staying current

The old scheduling feature helped—but it disappeared. Now we’re asking not just for its return, but for it to be rebuilt right.

* * *

**Requested Features:**

### 1. **Scheduling Options (Per URL):**

When uploading or managing a website Knowledge File:

- Add scrape frequency: **Manual** , **Daily** , **Weekly** , **Monthly**

- Allow creators to set **time of day** for scraping

- Optionally **retain or overwrite previous content** (with version tagging)

### 2. **Webhook Trigger Support (New):**

After a successful or failed scrape, Pickaxe should offer an **outbound webhook option** :

- Trigger a **custom URL** (e.g., n8n, Zapier, Make, or internal system)

- Include payload: file ID, timestamp, scrape status, and diff summary if applicable

- Supports automation like:

**Example Use Case:**  
Scrape site at 6am → webhook hits n8n → n8n notifies Slack + updates a Google Sheet + pings an AI workflow.

### 3. **After-Scrape Diff Reporting (Optional but Strongly Recommended):**

Generate and optionally send to Studio Owner/Controller:

- How many chunks were added/changed/removed

- Diff summary in plain text

- Timestamp + link to updated Knowledge File

### 4. **Scrape Health Monitoring:**

- Log scrape attempts with status (success, fail, skipped)

- Retry logic (e.g., 3x backoff)

- Email or in-app alert if a scrape fails repeatedly

* * *

**Why This Matters:**  
This feature turns static data into **living knowledge** , essential for agents supporting real-time or frequently updated domains. Adding webhook support also unlocks serious integration power for automation-minded users—without forcing them into brittle scraping workarounds outside Pickaxe.

* * *

**Final Thought:**  
Bring scraping back—but make it programmable, transparent, and reliable. Knowledge is only useful when it’s current. Let us keep it that way, on our own terms.

---

<div class="post-metadata">

**Author:** ![Ned.Malki](https://yyz2.discourse-cdn.com/flex004/user_avatar/community.pickaxe.co/ned.malki/32/5918_2.png) [@Ned.Malki](https://community.pickaxe.co/u/Ned.Malki)\
**Post date:** [July 31, 2025, 6:35am UTC](https://community.pickaxe.co/t/bring-back-and-improve-scheduled-scraping-for-knowledge-files-with-webhook-support/6733/2 "2025-07-31T06:35:20Z")

</div>

Hey @taedog2020, here’s more insight into why the scheduled scraping was rolled back in V2:

> [@How do I auto-update website content in my AI knowledge base?](https://community.pickaxe.co/t/how-do-i-auto-update-website-content-in-my-ai-knowledge-base/398/14):
>
> It was indeed a great feature, it just only worked 70% as well as it should due to some complications with the web-browsing and it created situatinos where users sometimes thought it was doing something it was not. Because of that, we rolled it back. If we release it again, it will work more efficiently.

---

<div class="post-metadata">

**Author:** ![taedog2020](https://yyz2.discourse-cdn.com/flex004/user_avatar/community.pickaxe.co/taedog2020/32/809_2.png) [@taedog2020](https://community.pickaxe.co/u/taedog2020)\
**Post date:** [August 1, 2025, 12:01pm UTC](https://community.pickaxe.co/t/bring-back-and-improve-scheduled-scraping-for-knowledge-files-with-webhook-support/6733/3 "2025-08-01T12:01:53Z")

</div>

That’s why I said add webhook. have n8n or something to do the scrape and update the KB.

---

<div class="post-metadata">

**Author:** ![Ned.Malki](https://yyz2.discourse-cdn.com/flex004/user_avatar/community.pickaxe.co/ned.malki/32/5918_2.png) [@Ned.Malki](https://community.pickaxe.co/u/Ned.Malki)\
**Post date:** [August 1, 2025, 12:03pm UTC](https://community.pickaxe.co/t/bring-back-and-improve-scheduled-scraping-for-knowledge-files-with-webhook-support/6733/4 "2025-08-01T12:03:46Z")

</div>

@taedog2020 You can do that now. Just add a webhook or connect an MCP server and configure your n8n or Make scenario. It’s already possible within Pickaxe.

---

<div class="post-metadata">

**Author:** ![taedog2020](https://yyz2.discourse-cdn.com/flex004/user_avatar/community.pickaxe.co/taedog2020/32/809_2.png) [@taedog2020](https://community.pickaxe.co/u/taedog2020)\
**Post date:** [August 1, 2025, 12:05pm UTC](https://community.pickaxe.co/t/bring-back-and-improve-scheduled-scraping-for-knowledge-files-with-webhook-support/6733/5 "2025-08-01T12:05:27Z")

</div>

so it will dynamically update the KB in the studio?

---

<div class="post-metadata">

**Author:** ![Ned.Malki](https://yyz2.discourse-cdn.com/flex004/user_avatar/community.pickaxe.co/ned.malki/32/5918_2.png) [@Ned.Malki](https://community.pickaxe.co/u/Ned.Malki)\
**Post date:** [August 1, 2025, 12:13pm UTC](https://community.pickaxe.co/t/bring-back-and-improve-scheduled-scraping-for-knowledge-files-with-webhook-support/6733/6 "2025-08-01T12:13:16Z")

</div>

That depends on your scenario. You can set up a scenario with a proxied scraper and set the refresh intervals that will trigger a website scrape.

**For example:**

 ![image](https://canada1.discourse-cdn.com/flex004/uploads/pickaxeproject/original/2X/9/99ec8798f95b361f3cd01b7348e98e6b53555ca6.png)
