Skip to content

  • Projects
  • Groups
  • Snippets
  • Help
    • Loading...
    • Help
    • Submit feedback
    • Contribute to GitLab
  • Sign in
U
upgrade-data-crawler-be
  • Project
    • Project
    • Details
    • Activity
    • Releases
    • Cycle Analytics
  • Repository
    • Repository
    • Files
    • Commits
    • Branches
    • Tags
    • Contributors
    • Graph
    • Compare
    • Charts
  • Issues 0
    • Issues 0
    • List
    • Board
    • Labels
    • Milestones
  • Merge Requests 0
    • Merge Requests 0
  • CI / CD
    • CI / CD
    • Pipelines
    • Jobs
    • Schedules
    • Charts
  • Wiki
    • Wiki
  • Snippets
    • Snippets
  • Members
    • Members
  • Collapse sidebar
  • Activity
  • Graph
  • Charts
  • Create a new issue
  • Jobs
  • Commits
  • Issue Boards
  • ThinhNC
  • upgrade-data-crawler-be
  • Merge Requests
  • !13

Merged
Opened Sep 07, 2026 by ThinhNC@ThinhNC
  • Report abuse
Report abuse

fix(crawl-jobs): implement soft delete for quota retention and auto-complete finished jobs

Summary of Changes

1. Daily Quota Retention via Soft Delete

  • Problem: Previously, deleting a crawl job triggered a hard delete (DELETE FROM crawl_jobs), causing countJobsSince and sumPagesCrawledByUser to decrease and prematurely resetting the user's daily quota (maxJobsPerDayLimit).
  • Solution:
    • Added deletedAt (DateTime?) and deletedBy (String?) to the CrawlJob model in schema.prisma with corresponding indexes (deleted_at, [user_id, deleted_at]).
    • Generated and applied migration 20260907123000_add_crawl_job_soft_delete.
    • Updated CrawlJobRepository.delete() to soft-delete the job record while cleaning up child tables (crawlAsset, crawlJobLog, crawlExport, crawlPage) and storage files.
    • Excluded soft-deleted jobs from findById, find, findByScheduleId, countConcurrentJobs, and dashboard stats queries.
    • Kept soft-deleted jobs included in countJobsSince and sumPagesCrawledByUser to preserve daily usage integrity.

2. Auto-completion for Stalled / Orphaned Finished Jobs

  • Problem: Seed job 088f635c-9c3a-4467-93bb-e58f001bf001 (and potential stalled jobs) remained in RUNNING status indefinitely even though all 52 target pages were already processed (49 success, 3 failed), preventing data export.
  • Solution:
    • In CrawlJobService.findById and findAllByUser, automatically detect jobs in RUNNING status where processedPages >= targetPages and update them to COMPLETED with finishedAt.
    • In CrawlExportService.createExport, allow export generation by auto-completing finished jobs before validating export readiness.
    • Updated prisma/seed.ts to set sample Job 1 status to COMPLETED.

Check out, review, and merge locally

Step 1. Fetch and check out the branch for this merge request

git fetch origin
git checkout -b fix/soft-delete-quota-retention-and-job-completion origin/fix/soft-delete-quota-retention-and-job-completion

Step 2. Review the changes locally

Step 3. Merge the branch and fix any conflicts that come up

git fetch origin
git checkout origin/develop
git merge --no-ff fix/soft-delete-quota-retention-and-job-completion

Step 4. Push the result of the merge to GitLab

git push origin develop

Note that pushing to GitLab requires write access to this repository.

Tip: You can also checkout merge requests locally by following these guidelines.

  • Discussion 0
  • Commits 1
  • Changes 10
Assignee
No assignee
Assign to
None
Milestone
None
Assign milestone
Time tracking
0
Labels
None
Assign labels
  • View project labels
Reference: ThinhNC/upgrade-data-crawler-be!13

Revert this merge request

This will create a new commit in order to revert the existing changes.

Switch branch
Cancel
A new branch will be created in your fork and a new merge request will be started.

Cherry-pick this merge request

Switch branch
Cancel
A new branch will be created in your fork and a new merge request will be started.