Commit 3f53534f authored by ThinhNC's avatar ThinhNC

Merge branch 'feat/export-zip-bundles-and-data-contract' into 'develop'

Feat/export zip bundles and data contract

See merge request !24
parents 398fe4df 279b122a
......@@ -338,15 +338,37 @@ Khi xuất file gói ZIP, hệ thống phân chia rõ ràng thư mục clean/raw
export-job-c4b8e21a.zip
├── data/
│ ├── raw/pages.raw.json # Dữ liệu JSON thô (chứa rawMarkdown)
│ └── clean/pages.clean.json # Dữ liệu JSON sạch (chứa mainContent & cleanText)
│ ├── clean/pages.clean.json # Dữ liệu JSON sạch (chứa mainContent & cleanText)
│ ├── pages.json # Envelope JSON tổng thể
│ ├── pages.csv # Bảng CSV danh sách trang kèm chỉ số chất lượng & pageId
│ ├── links.csv # Danh sách liên kết nội/ngoại bộ (CSV)
│ ├── images.csv # Danh sách hình ảnh (CSV)
│ ├── pages.xlsx # Bảng tính Excel danh sách trang
│ └── tables.xlsx # Dữ liệu các bảng HTML dưới dạng Excel
├── markdown/
│ ├── raw/*.md # Các file .md thô nguyên bản
│ └── clean/*.md # Các file .md sạch đã lọc bỏ nav/footer
├── tables.xlsx # Dữ liệu các bảng HTML dưới dạng Excel
├── links.csv # Danh sách liên kết nội/ngoại bộ
├── images.csv # Danh sách thông tin hình ảnh
├── logs/
│ ├── errors.json # Báo cáo các trang bị lỗi
│ └── crawl-log.txt # Nhật ký chi tiết tiến trình crawl
├── metadata.json # Tổng quan thông số job
└── errors.json # Báo cáo các trang bị lỗi
├── summary.json # Thống kê tổng hợp số lượng trang
├── data_quality.json # Báo cáo điểm chất lượng & cảnh báo
└── diff_report.json # Báo cáo thay đổi nội dung (Change Detection)
```
### 8.5 Các Gói Xuất Riêng Lẻ (Individual Export Formats)
Ngoài gói ZIP tổng thể (`exportType: "ZIP"`), hệ thống hỗ trợ xuất độc lập từng định dạng:
- **`exportType: "XLSX"`**: Tự động đóng gói thành **`xlsx.zip`** chứa:
- `pages.xlsx`: Danh sách trang với Page ID, Word Count, Data Quality Score, Content Preview.
- `tables.xlsx`: Trích xuất toàn bộ bảng HTML thành các Sheet kèm trang `Summary` (có Page ID đối chiếu).
- **`exportType: "CSV"`**: Tự động đóng gói thành **`csv.zip`** chứa: `pages.csv`, `links.csv`, `images.csv`.
- **`exportType: "MARKDOWN"`**: Tự động đóng gói thành **`markdown.zip`** chứa: `clean/*.md``raw/*.md`.
- **`exportType: "JSON"`**: Tự động đóng gói thành **`json.zip`** chứa:
- `pages.json`: Master envelope đầy đủ nhất theo chuẩn Data Contract v1.
- `clean/pages.clean.json`: Dữ liệu sạch cho AI/LLM Prompt Context.
- `raw/pages.raw.json`: Dữ liệu thô nguyên bản phục vụ audit.
- `structured.json`: Dữ liệu có cấu trúc schema.org (nếu có).
Xem chi tiết Data Contract đầy đủ tại [DATA_CONTRACT_V1.md](docs/DATA_CONTRACT_V1.md).
......@@ -201,81 +201,88 @@ Khi người dùng khởi tạo yêu cầu xuất dữ liệu (`POST /crawl-jobs
Chứa Envelope tổng thể như đã trình bày ở Mục 2.
### 4.2 `links.csv`
### 4.2 `pages.csv`
File bảng CSV tổng hợp danh sách các trang đã thu thập, tối ưu cho việc mở xem nhanh bằng Excel / Google Sheets hoặc nhập liệu vào database:
- **Headers**: `pageId,url,title,description,status,statusCode,errorMessage,wordCount,dataQualityScore,mainContent,crawledAt`
- Cột `pageId` là khóa ngoại (foreign key) map trực tiếp với `links.csv``images.csv`.
- Nội dung văn bản chỉ giữ trường sạch `mainContent` để tránh phình dung lượng và tránh lỗi giới hạn ký tự ô của Excel (32,767 ký tự).
### 4.3 `links.csv`
Chứa tất cả các liên kết thu thập được từ toàn bộ các trang trong job.
- **Headers**: `pageId,sourceUrl,url,type`
### 4.3 `images.csv`
### 4.4 `images.csv`
Chứa thông tin tất cả hình ảnh thu thập được trong job.
- **Headers**: `pageId,sourceUrl,altText,orderIndex,type`
### 4.4 `tables.xlsx`
### 4.5 `tables.xlsx` & `pages.xlsx`
File bảng tính Excel tổng hợp toàn bộ các bảng HTML được phát hiện trong các trang:
- **`tables.xlsx`**: File bảng tính Excel tổng hợp toàn bộ các bảng HTML được phát hiện trong các trang (Sheet "Summary" và các sheet chi tiết của từng bảng).
- **`pages.xlsx`**: File bảng tính Excel dành cho người dùng xem nhanh danh sách các trang đã crawl kèm định dạng màu trạng thái.
- **Sheet "Summary"**: Tổng hợp danh sách các bảng, URL trang chứa, số hàng, số cột và tên worksheet chi tiết.
- **Các Sheet "Table-P<ShortPath>-<Index>"**: Chứa dữ liệu chi tiết dạng lưới ô (cells) của từng bảng.
### 4.6 `metadata.json`, `summary.json`, `data_quality.json`, `diff_report.json`
### 4.5 `pages.xlsx`
- **`metadata.json`**: Tóm tắt tổng quan thông số và cấu hình chạy của Crawl Job.
- **`summary.json`**: Thống kê số lượng trang thành công/thất bại và thời gian hoàn thành.
- **`data_quality.json`**: Báo cáo tổng hợp điểm chất lượng dữ liệu, tỷ lệ nội dung sạch, các cảnh báo (nav noise, trùng lặp, bài viết quá ngắn).
- **`diff_report.json`**: Báo cáo phát hiện thay đổi nội dung (Change Detection) giữa các lần crawl.
File Excel bảng tính dành cho người dùng xem nhanh danh sách các trang đã crawl, trạng thái, mã HTTP, tiêu đề, mô tả, điểm chất lượng và số từ.
### 4.7 `logs/errors.json` & `logs/crawl-log.txt`
### 4.6 `metadata.json`
- **`errors.json`**: Báo cáo riêng các trang bị lỗi (`status !== 'SUCCESS'`) kèm mã lỗi và nguyên nhân chi tiết.
- **`crawl-log.txt`**: Toàn bộ nhật ký chạy tiến trình crawl.
Tóm tắt tổng quan tiến trình chạy của Crawl Job:
### 4.8 Gói Xuất CSV Riêng Lẻ (`exportType: "CSV"`)
```json
{
"jobId": "c4b8e21a-4d3f-4e89-9a1b-2c3d4e5f6a7b",
"userId": "9f8e7d6c-5b4a-3f2e-1d0c-9b8a7f6e5d4c",
"startUrl": "https://example.com",
"domain": "example.com",
"mode": "CRAWL",
"status": "COMPLETED",
"maxPages": 100,
"maxDepth": 3,
"totalPages": 50,
"successPages": 48,
"failedPages": 2,
"timeoutMs": 30000,
"retryCount": 3,
"startedAt": "2026-07-21T13:20:00.000Z",
"finishedAt": "2026-07-21T13:28:00.000Z",
"exportedAt": "2026-07-21T13:30:00.000Z"
}
```
Khi chọn xuất định dạng `CSV`, hệ thống tự động đóng gói toàn bộ các bảng CSV thành tệp **`csv.zip`** chứa:
- `pages.csv`: Bảng tổng hợp trang kèm chỉ số chất lượng và ID.
- `links.csv`: Bảng liên kết nội/ngoại bộ.
- `images.csv`: Bảng danh sách hình ảnh trích xuất.
### 4.7 `errors.json`
### 4.9 Gói Xuất Markdown Riêng Lẻ (`exportType: "MARKDOWN"`)
Báo cáo riêng các trang bị lỗi (`status !== 'SUCCESS'`):
Hệ thống đóng gói toàn bộ tài liệu Markdown thành **`markdown.zip`** gồm 2 thư mục:
- `clean/*.md`: File Markdown sạch đã bóc tách nav/footer dành cho AI Prompt Context.
- `raw/*.md`: File Markdown thô nguyên bản phục vụ kiểm tra/đối chiếu.
```json
[
{
"url": "https://example.com/protected-page",
"status": "REQUIRES_LOGIN",
"statusCode": 401,
"errorMessage": "Page requires authentication credentials",
"crawledAt": "2026-07-21T13:26:00.000Z"
}
]
```
### 4.10 Gói Xuất XLSX Riêng Lẻ (`exportType: "XLSX"`)
Khi chọn xuất định dạng `XLSX`, hệ thống tự động đóng gói toàn bộ bảng tính Excel vào tệp **`xlsx.zip`** chứa:
- `pages.xlsx`: Danh sách toàn bộ các trang crawl kèm Page ID, Word Count, Data Quality Score, Content Preview và metadata.
- `tables.xlsx`: Tập hợp toàn bộ bảng HTML trích xuất được từ website, bao gồm trang `Summary` (liệt kê danh sách bảng kèm Page ID và URL để đối chiếu chéo) và từng Sheet cho từng bảng dữ liệu riêng biệt.
### 4.8 Cấu trúc File ZIP Export (`exportType: "ZIP"`)
### 4.11 Gói Xuất JSON Riêng Lẻ (`exportType: "JSON"`)
Khi chọn export định dạng `ZIP`, gói lưu trữ sẽ tự động cấu trúc phân cấp dữ liệu clean/raw thành các thư mục riêng biệt:
Khi chọn xuất định dạng `JSON`, hệ thống tự động đóng gói toàn bộ các file JSON dữ liệu vào tệp **`json.zip`** chứa:
- `pages.json`: Master Envelope đầy đủ nhất theo chuẩn Data Contract v1.
- `clean/pages.clean.json`: Dữ liệu sạch đã lọc bỏ `rawMarkdown`, tối ưu hóa token cho AI Prompt Context và RAG Indexing.
- `raw/pages.raw.json`: Dữ liệu thô nguyên bản phục vụ audit / debug.
- `structured.json`: Dữ liệu có cấu trúc (schema.org JSON-LD / OpenGraph), chỉ xuất hiện khi có ít nhất một trang có dữ liệu này.
### 4.12 Cấu trúc File ZIP Xuất Toàn Bộ (`exportType: "ZIP"`)
Khi chọn export định dạng `ZIP`, gói lưu trữ chứa toàn bộ dữ liệu phân cấp theo đúng cấu trúc tiêu chuẩn:
```
export-job-c4b8e21a.zip
├── data/
│ ├── raw/
│ │ └── pages.raw.json # Danh sách trang thô (chứa rawMarkdown)
│ └── clean/
│ └── pages.clean.json # Danh sách trang sạch (chứa mainContent & cleanText)
│ ├── clean/
│ │ └── pages.clean.json # Danh sách trang sạch (chứa mainContent & cleanText)
│ ├── pages.json # Dữ liệu Envelope đầy đủ
│ ├── structured.json # Dữ liệu trích xuất có cấu trúc
│ ├── pages.csv # Bảng CSV danh sách trang kèm chỉ số chất lượng
│ ├── links.csv # Danh sách liên kết nội/ngoại bộ (CSV)
│ ├── images.csv # Danh sách hình ảnh (CSV)
│ ├── pages.xlsx # Bảng tính Excel danh sách trang
│ └── tables.xlsx # Bảng tính Excel chi tiết các HTML tables
├── markdown/
│ ├── raw/
│ │ ├── page-1.md # Tệp markdown thô nguyên bản của từng trang
......@@ -283,11 +290,13 @@ export-job-c4b8e21a.zip
│ └── clean/
│ ├── page-1.md # Tệp markdown sạch đã lọc nav/footer
│ └── page-2.md
├── tables.xlsx # Bảng tính chứa dữ liệu các HTML tables
├── links.csv # Danh sách liên kết nội/ngoại bộ
├── images.csv # Danh sách thông tin hình ảnh
├── logs/
│ ├── errors.json # Nhật ký lỗi các trang thất bại
│ └── crawl-log.txt # Nhật ký tiến trình crawl
├── metadata.json # Tổng quan thông số job
└── errors.json # Nhật ký lỗi trang thất bại
├── summary.json # Thống kê tổng hợp kết quả
├── data_quality.json # Báo cáo điểm chất lượng & cảnh báo
└── diff_report.json # Báo cáo thay đổi nội dung (nếu có)
```
---
......
export const parse = () => ({
querySelectorAll: () => [],
export interface MockNode {
tagName: string;
text: string;
rawText: string;
querySelector: (selector: string) => MockNode | null;
querySelectorAll: (selector: string) => MockNode[];
}
function stripTags(html: string): string {
return html.replace(/<[^>]*>/g, "").trim();
}
function parseElement(tag: string, content: string): MockNode {
const innerText = stripTags(content);
return {
tagName: tag.toLowerCase(),
text: innerText,
rawText: innerText,
querySelector(selector: string) {
const results = this.querySelectorAll(selector);
return results.length > 0 ? results[0] : null;
},
querySelectorAll(selector: string) {
return querySelectorAllFromHtml(content, selector);
},
};
}
function querySelectorAllFromHtml(html: string, selector: string): MockNode[] {
const selectors = selector.split(",").map((s) => s.trim());
const found: MockNode[] = [];
for (const sel of selectors) {
let targetTag = sel.toLowerCase();
if (targetTag.includes(" ")) {
const parts = targetTag.split(/\s+/);
targetTag = parts[parts.length - 1];
}
targetTag = targetTag.replace(/:[a-zA-Z0-9_-]+(\([^)]*\))?/g, "");
const regex = new RegExp(
`<${targetTag}\\b[^>]*>([\\s\\S]*?)<\\/${targetTag}>`,
"gi",
);
let match: RegExpExecArray | null;
while ((match = regex.exec(html)) !== null) {
found.push(parseElement(targetTag, match[1]));
}
}
return found;
}
export const parse = (html: string = "") => ({
querySelector: (selector: string) => {
const list = querySelectorAllFromHtml(html, selector);
return list.length > 0 ? list[0] : null;
},
querySelectorAll: (selector: string) => querySelectorAllFromHtml(html, selector),
});
......@@ -9,9 +9,12 @@ export const EXPORT_TYPE = {
export type ExportType = keyof typeof EXPORT_TYPE;
export const EXPORT_MIME_TYPES: Record<ExportType, string> = {
JSON: "application/json",
CSV: "text/csv",
XLSX: "application/vnd.openxmlformats-officedocument.spreadsheetml.sheet",
// JSON export tạo file .zip chứa các file JSON (pages.json, clean/pages.clean.json, raw/pages.raw.json, structured.json)
JSON: "application/zip",
// CSV export tạo file .zip chứa các file CSV (pages.csv, links.csv, images.csv)
CSV: "application/zip",
// XLSX export tạo file .zip chứa các file Excel (pages.xlsx, tables.xlsx)
XLSX: "application/zip",
// MARKDOWN export tạo file .zip chứa nhiều file markdown — nên dùng application/zip
MARKDOWN: "application/zip",
ZIP: "application/zip",
......
......@@ -42,3 +42,15 @@ export function buildCrawlResultZipKey(jobId: string): string {
export function buildMarkdownZipName(jobId: string): string {
return "markdown.zip";
}
export function buildCsvZipName(jobId: string): string {
return "csv.zip";
}
export function buildXlsxZipName(jobId: string): string {
return "xlsx.zip";
}
export function buildJsonZipName(jobId: string): string {
return "json.zip";
}
......@@ -5,6 +5,9 @@ import {
JOB_EXPORT_SUBDIRS,
buildCrawlResultZipName,
buildMarkdownZipName,
buildCsvZipName,
buildXlsxZipName,
buildJsonZipName,
} from "../constants/storage-path.constant";
import { generatePageFileName } from "./slug.helper";
......@@ -142,6 +145,33 @@ export function buildJobMarkdownZipPath(jobId: string): {
return { fileName, filePath: path.join(rootDir, fileName) };
}
export function buildJobCsvZipPath(jobId: string): {
fileName: string;
filePath: string;
} {
const fileName = buildCsvZipName(jobId);
const rootDir = buildJobRootDir(jobId);
return { fileName, filePath: path.join(rootDir, fileName) };
}
export function buildJobXlsxZipPath(jobId: string): {
fileName: string;
filePath: string;
} {
const fileName = buildXlsxZipName(jobId);
const rootDir = buildJobRootDir(jobId);
return { fileName, filePath: path.join(rootDir, fileName) };
}
export function buildJobJsonZipPath(jobId: string): {
fileName: string;
filePath: string;
} {
const fileName = buildJsonZipName(jobId);
const rootDir = buildJobRootDir(jobId);
return { fileName, filePath: path.join(rootDir, fileName) };
}
export function buildExportFilePath(
exportDir: string,
fileName: string,
......
import fs from "fs";
import { EventEmitter } from "events";
import { CsvExportService } from "../csv-export.service";
jest.mock("../../../database/prisma.client", () => ({
......@@ -7,7 +8,24 @@ jest.mock("../../../database/prisma.client", () => ({
jest.mock("../../crawl-assets/crawl-asset.repository", () => ({
CrawlAssetRepository: jest.fn().mockImplementation(() => ({
findByJobId: jest.fn().mockResolvedValue([]),
findByJobId: jest.fn().mockResolvedValue([
{
id: "asset-1",
pageId: "page-1",
assetType: "LINK",
url: "https://example.com/sub",
sourceUrl: "https://example.com/test",
},
{
id: "asset-2",
pageId: "page-1",
assetType: "IMAGE",
url: "https://example.com/image.png",
altText: "Test Image",
mimeType: "image/png",
orderIndex: 1,
},
]),
})),
}));
......@@ -16,9 +34,30 @@ jest.mock("../../../common/helpers/file.helper", () => ({
fileName,
filePath: `test/${fileName}`,
})),
buildJobCsvZipPath: jest.fn((jobId: string) => ({
fileName: "csv.zip",
filePath: `test/csv.zip`,
})),
ensureJobExportStructure: jest.fn(),
ensureDirExists: jest.fn(),
getFileSizeBytes: jest.fn(() => 1024),
}));
let mockArchiveStream: EventEmitter;
jest.mock("archiver", () => {
return jest.fn(() => ({
pipe: jest.fn(),
file: jest.fn(),
finalize: jest.fn().mockImplementation(function (this: any) {
if (mockArchiveStream) {
process.nextTick(() => mockArchiveStream.emit("close"));
}
}),
on: jest.fn(),
}));
});
jest.mock("fs");
describe("CsvExportService", () => {
......@@ -27,9 +66,13 @@ describe("CsvExportService", () => {
beforeEach(() => {
jest.clearAllMocks();
service = new CsvExportService();
mockArchiveStream = new EventEmitter();
(fs.createWriteStream as jest.Mock).mockImplementation(() => mockArchiveStream);
(fs.existsSync as jest.Mock).mockReturnValue(true);
});
it("neutralizes formula injection characters (=, +, -, @) in CSV export", async () => {
it("neutralizes formula injection characters (=, +, -, @) and includes pageId in pages.csv", async () => {
const writtenFiles: Record<string, string> = {};
(fs.writeFileSync as jest.Mock).mockImplementation((filePath, content) => {
writtenFiles[filePath] = content;
......@@ -47,18 +90,67 @@ describe("CsvExportService", () => {
description: "@SUM(1,2)",
status: "COMPLETED",
statusCode: 200,
errorMessage: null,
wordCount: 150,
dataQualityScore: 90,
markdownContent: "+12345",
crawledAt: new Date("2026-09-02T12:00:00Z"),
},
],
};
await (service as any).executeExport(mockJob);
await service.exportPages(mockJob);
const pagesCsv = writtenFiles["test/pages.csv"];
expect(pagesCsv).toBeDefined();
// Verify headers
const headerLine = pagesCsv.split("\n")[0];
expect(headerLine).toBe(
"pageId,url,title,description,status,statusCode,errorMessage,wordCount,dataQualityScore,mainContent,crawledAt",
);
// Verify content & formula neutralization
expect(pagesCsv).toContain("page-1");
expect(pagesCsv).toContain("'=cmd|'/C calc'!A0");
expect(pagesCsv).toContain("'@SUM(1,2)");
expect(pagesCsv).toContain("'+12345");
});
it("exports pages.csv, links.csv, images.csv and bundles them into csv.zip", async () => {
const writtenFiles: Record<string, string> = {};
(fs.writeFileSync as jest.Mock).mockImplementation((filePath, content) => {
writtenFiles[filePath] = content;
});
const mockJob: any = {
id: "job-1",
startUrl: "https://example.com",
domain: "example.com",
pages: [
{
id: "page-1",
url: "https://example.com/test",
title: "Test Page",
description: "Description",
status: "COMPLETED",
statusCode: 200,
errorMessage: null,
wordCount: 100,
dataQualityScore: 85,
markdownContent: "Hello world",
crawledAt: new Date("2026-09-02T12:00:00Z"),
},
],
};
const result = await service.export(mockJob);
expect(writtenFiles["test/pages.csv"]).toBeDefined();
expect(writtenFiles["test/links.csv"]).toBeDefined();
expect(writtenFiles["test/images.csv"]).toBeDefined();
expect(result.fileName).toBe("csv.zip");
expect(result.filePath).toContain("csv.zip");
});
});
import fs from "fs";
import { EventEmitter } from "events";
import { JsonExportService } from "../json-export.service";
import { JOB_EXPORT_FILES } from "../../../common/constants/storage-path.constant";
jest.mock("../../../database/prisma.client", () => ({
prisma: {},
}));
jest.mock("../../crawl-assets/crawl-asset.repository", () => ({
CrawlAssetRepository: jest.fn().mockImplementation(() => ({
findAssetsForJsonExport: jest.fn().mockResolvedValue([
{
id: "asset-1",
pageId: "page-1",
assetType: "LINK",
url: "https://example.com/sub",
sourceUrl: "https://example.com/test",
},
{
id: "asset-2",
pageId: "page-1",
assetType: "IMAGE",
url: "https://example.com/image.png",
altText: "Test Image",
mimeType: "image/png",
orderIndex: 1,
},
]),
})),
}));
jest.mock("../../../common/helpers/file.helper", () => ({
buildJobDataFilePath: jest.fn((jobId: string, fileName: string) => ({
fileName,
filePath: `test/data/${fileName}`,
})),
buildJobDataRawFilePath: jest.fn((jobId: string, fileName: string) => ({
fileName,
filePath: `test/data/raw/${fileName}`,
})),
buildJobDataCleanFilePath: jest.fn((jobId: string, fileName: string) => ({
fileName,
filePath: `test/data/clean/${fileName}`,
})),
buildJobJsonZipPath: jest.fn((jobId: string) => ({
fileName: "json.zip",
filePath: `test/json.zip`,
})),
ensureJobExportStructure: jest.fn(),
ensureDirExists: jest.fn(),
getFileSizeBytes: jest.fn(() => 1024),
}));
let mockArchiveStream: EventEmitter;
let archivedFiles: { path: string; name: string }[] = [];
jest.mock("archiver", () => {
return jest.fn(() => ({
pipe: jest.fn(),
file: jest.fn((sourcePath: string, data: { name: string }) => {
archivedFiles.push({ path: sourcePath, name: data.name });
}),
finalize: jest.fn().mockImplementation(function (this: any) {
if (mockArchiveStream) {
process.nextTick(() => mockArchiveStream.emit("close"));
}
}),
on: jest.fn(),
}));
});
describe("JsonExportService", () => {
let service: JsonExportService;
let writtenFiles: Record<string, string> = {};
beforeEach(() => {
jest.clearAllMocks();
writtenFiles = {};
archivedFiles = [];
service = new JsonExportService();
mockArchiveStream = new EventEmitter();
jest.spyOn(fs, "createWriteStream").mockImplementation(() => mockArchiveStream as any);
jest.spyOn(fs, "existsSync").mockReturnValue(true);
jest.spyOn(fs, "writeFileSync").mockImplementation((filePath, data) => {
writtenFiles[filePath.toString()] = data.toString();
});
});
it("exports pages.json, clean/pages.clean.json, raw/pages.raw.json and bundles into json.zip", async () => {
const mockJob: any = {
id: "job-1",
startUrl: "https://example.com",
domain: "example.com",
pages: [
{
id: "page-1",
jobId: "job-1",
url: "https://example.com/test",
content: "<div>Content with <table><tr><td>Item</td></tr></table></div>",
markdownContent: "# Title\n\nMain content for test.\n\n[Link](https://example.com/sub)",
title: "Test Page",
description: "Test Description",
status: "SUCCESS",
statusCode: 200,
crawledAt: new Date("2026-07-21T10:00:00.000Z"),
structuredData: null,
},
],
};
const result = await service.export(mockJob);
expect(result.fileName).toBe("json.zip");
expect(result.filePath).toMatch(/test[\\/]json\.zip$/);
// 1. Kiểm tra pages.json
const pagesJsonContent = JSON.parse(writtenFiles[`test/data/${JOB_EXPORT_FILES.PAGES_JSON}`]);
expect(pagesJsonContent.jobId).toBe("job-1");
expect(pagesJsonContent.schemaVersion).toBeDefined();
expect(pagesJsonContent.totalRecords).toBe(1);
expect(pagesJsonContent.pages[0].id).toBe("page-1");
expect(pagesJsonContent.pages[0].rawMarkdown).toBeDefined();
expect(pagesJsonContent.pages[0].cleanText).toBeDefined();
expect(pagesJsonContent.pages[0].mainContent).toBeDefined();
// 2. Kiểm tra pages.clean.json: không có rawMarkdown
const cleanJsonContent = JSON.parse(writtenFiles[`test/data/clean/${JOB_EXPORT_FILES.PAGES_CLEAN_JSON}`]);
expect(cleanJsonContent.pages[0].rawMarkdown).toBeUndefined();
expect(cleanJsonContent.pages[0].mainContent).toBeDefined();
expect(cleanJsonContent.pages[0].cleanText).toBeDefined();
expect(cleanJsonContent.pages[0].dataQualityScore).toBeDefined();
// 3. Kiểm tra pages.raw.json: không có cleanText và dataQualityScore
const rawJsonContent = JSON.parse(writtenFiles[`test/data/raw/${JOB_EXPORT_FILES.PAGES_RAW_JSON}`]);
expect(rawJsonContent.pages[0].rawMarkdown).toBeDefined();
expect(rawJsonContent.pages[0].cleanText).toBeUndefined();
expect(rawJsonContent.pages[0].dataQualityScore).toBeUndefined();
// 4. Kiểm tra các file được đưa vào zip
const archivedNames = archivedFiles.map((f) => f.name);
expect(archivedNames).toContain("pages.json");
expect(archivedNames).toContain("clean/pages.clean.json");
expect(archivedNames).toContain("raw/pages.raw.json");
});
it("writes structured.json and packs it when pages have structuredData", async () => {
const mockJob: any = {
id: "job-2",
startUrl: "https://example.com",
domain: "example.com",
pages: [
{
id: "page-1",
jobId: "job-2",
url: "https://example.com/item",
content: "<div>Content</div>",
markdownContent: "Content",
status: "SUCCESS",
statusCode: 200,
structuredData: { "@type": "Product", name: "Widget" },
},
],
};
await service.export(mockJob);
const structuredJsonPath = `test/data/${JOB_EXPORT_FILES.STRUCTURED_JSON}`;
expect(writtenFiles[structuredJsonPath]).toBeDefined();
const structuredContent = JSON.parse(writtenFiles[structuredJsonPath]);
expect(structuredContent.records).toHaveLength(1);
expect(structuredContent.records[0].structuredData).toEqual({ "@type": "Product", name: "Widget" });
const archivedNames = archivedFiles.map((f) => f.name);
expect(archivedNames).toContain("structured.json");
});
it("does not write structured.json when no pages have structuredData", async () => {
const mockJob: any = {
id: "job-3",
startUrl: "https://example.com",
domain: "example.com",
pages: [
{
id: "page-1",
jobId: "job-3",
url: "https://example.com/no-data",
content: "<div>Content</div>",
markdownContent: "Content",
status: "SUCCESS",
statusCode: 200,
structuredData: null,
},
],
};
// Khi structured.json không tồn tại, existsSync trả về false
jest.spyOn(fs, "existsSync").mockImplementation((p: any) => {
if (p.toString().includes("structured.json")) return false;
return true;
});
await service.export(mockJob);
const structuredJsonPath = `test/data/${JOB_EXPORT_FILES.STRUCTURED_JSON}`;
expect(writtenFiles[structuredJsonPath]).toBeUndefined();
const archivedNames = archivedFiles.map((f) => f.name);
expect(archivedNames).not.toContain("structured.json");
});
});
import fs from "fs";
import { MarkdownExportService } from "../markdown-export.service";
jest.mock("../../../database/prisma.client", () => ({
prisma: {},
}));
jest.mock("../../../common/helpers/file.helper", () => ({
buildJobMarkdownRawFilePath: jest.fn((_jobId: string, idx: number) => ({
fileName: `00${idx}-raw.md`,
filePath: `test/markdown/raw/00${idx}-raw.md`,
})),
buildJobMarkdownCleanFilePath: jest.fn((_jobId: string, idx: number) => ({
fileName: `00${idx}-clean.md`,
filePath: `test/markdown/clean/00${idx}-clean.md`,
})),
buildJobMarkdownZipPath: jest.fn(() => ({
fileName: "markdown.zip",
filePath: "test/markdown.zip",
})),
buildJobSubDir: jest.fn(() => "test/markdown"),
ensureJobExportStructure: jest.fn(),
}));
jest.mock("fs");
describe("MarkdownExportService", () => {
let service: MarkdownExportService;
beforeEach(() => {
jest.clearAllMocks();
service = new MarkdownExportService();
});
it("writes only to raw/ and clean/ subdirectories without redundant files in root markdown/", () => {
const writtenFiles: Record<string, string> = {};
(fs.writeFileSync as jest.Mock).mockImplementation((filePath, content) => {
writtenFiles[filePath] = content;
});
const mockJob: any = {
id: "job-1",
startUrl: "https://example.com",
domain: "example.com",
pages: [
{
id: "page-1",
url: "https://example.com/test",
title: "Test Page",
markdownContent: "# Test Header\n\nMain content here.",
},
],
};
const results = service.writePageFiles(mockJob);
// Exactly 2 files should be written per page: one in raw, one in clean
expect(results).toHaveLength(2);
expect(Object.keys(writtenFiles)).toEqual([
"test/markdown/raw/000-raw.md",
"test/markdown/clean/000-clean.md",
]);
// Raw file contains original markdown
expect(writtenFiles["test/markdown/raw/000-raw.md"]).toContain("# Test Header");
// Clean file contains extracted content
expect(writtenFiles["test/markdown/clean/000-clean.md"]).toContain("# Test Header");
});
});
import fs from "fs";
import { EventEmitter } from "events";
import ExcelJS from "exceljs";
import { XlsxExportService } from "../xlsx-export.service";
jest.mock("../../../database/prisma.client", () => ({
prisma: {},
}));
jest.mock("../../../common/helpers/file.helper", () => ({
buildJobDataFilePath: jest.fn((jobId: string, fileName: string) => ({
fileName,
filePath: `test/${fileName}`,
})),
buildJobXlsxZipPath: jest.fn((jobId: string) => ({
fileName: "xlsx.zip",
filePath: `test/xlsx.zip`,
})),
ensureJobExportStructure: jest.fn(),
ensureDirExists: jest.fn(),
getFileSizeBytes: jest.fn(() => 2048),
}));
let mockArchiveStream: EventEmitter;
jest.mock("archiver", () => {
return jest.fn(() => ({
pipe: jest.fn(),
file: jest.fn(),
finalize: jest.fn().mockImplementation(function (this: any) {
if (mockArchiveStream) {
process.nextTick(() => mockArchiveStream.emit("close"));
}
}),
on: jest.fn(),
}));
});
describe("XlsxExportService", () => {
let service: XlsxExportService;
beforeEach(() => {
jest.clearAllMocks();
service = new XlsxExportService();
mockArchiveStream = new EventEmitter();
jest.spyOn(fs, "createWriteStream").mockImplementation(() => mockArchiveStream as any);
jest.spyOn(fs, "existsSync").mockReturnValue(true);
});
afterEach(() => {
jest.restoreAllMocks();
});
it("exports pages.xlsx with Page ID, Word Count, Quality Score, and Content Preview", async () => {
const addWorksheetSpy = jest.spyOn(ExcelJS.Workbook.prototype, "addWorksheet");
jest.spyOn(ExcelJS.Workbook.prototype.xlsx, "writeFile").mockResolvedValue(undefined as any);
const mockJob: any = {
id: "job-1",
startUrl: "https://example.com",
domain: "example.com",
pages: [
{
id: "page-1",
url: "https://example.com/test",
title: "Test Page",
description: "Test Description",
status: "COMPLETED",
statusCode: 200,
errorMessage: null,
wordCount: 250,
dataQualityScore: 95,
markdownContent: "# Header\n\nMain content here.",
crawledAt: new Date("2026-09-02T12:00:00Z"),
},
],
};
await service.exportPages(mockJob);
expect(addWorksheetSpy).toHaveBeenCalled();
const sheet = addWorksheetSpy.mock.results[0].value as ExcelJS.Worksheet;
expect(sheet.name).toBe("Pages");
const columnHeaders = sheet.columns?.map((c) => c.header) ?? [];
expect(columnHeaders).toContain("Page ID");
expect(columnHeaders).toContain("Word Count");
expect(columnHeaders).toContain("Quality Score");
expect(columnHeaders).toContain("Content Preview");
const dataRow = sheet.getRow(2).values as any[];
expect(dataRow).toContain("page-1");
expect(dataRow).toContain(250);
expect(dataRow).toContain(95);
});
it("exports tables.xlsx with Page ID in Summary sheet", async () => {
const addWorksheetSpy = jest.spyOn(ExcelJS.Workbook.prototype, "addWorksheet");
jest.spyOn(ExcelJS.Workbook.prototype.xlsx, "writeFile").mockResolvedValue(undefined as any);
const mockJob: any = {
id: "job-1",
startUrl: "https://example.com",
domain: "example.com",
pages: [
{
id: "page-1",
url: "https://example.com/test",
content: `
<div>
<table>
<caption>Pricing Table</caption>
<thead><tr><th>Item</th><th>Price</th></tr></thead>
<tbody><tr><td>Widget</td><td>$10</td></tr></tbody>
</table>
</div>
`,
markdownContent: "",
status: "COMPLETED",
},
],
};
await service.exportTables(mockJob);
expect(addWorksheetSpy).toHaveBeenCalled();
const summarySheet = addWorksheetSpy.mock.results.find(
(r) => r.value.name === "Summary",
)?.value as ExcelJS.Worksheet;
expect(summarySheet).toBeDefined();
const columnHeaders = summarySheet.columns?.map((c) => c.header) ?? [];
expect(columnHeaders).toContain("Page ID");
expect(columnHeaders).toContain("Sheet Name");
const dataRow = summarySheet.getRow(2).values as any[];
expect(dataRow).toContain("page-1");
});
it("bundles pages.xlsx and tables.xlsx into xlsx.zip upon export", async () => {
jest.spyOn(ExcelJS.Workbook.prototype.xlsx, "writeFile").mockResolvedValue(undefined as any);
const mockJob: any = {
id: "job-1",
startUrl: "https://example.com",
domain: "example.com",
pages: [],
};
const result = await service.export(mockJob);
expect(result.fileName).toBe("xlsx.zip");
expect(result.filePath).toContain("xlsx.zip");
});
});
import fs from "fs";
import archiver from "archiver";
import {
CrawlAsset,
CrawlJob,
......@@ -9,14 +10,12 @@ import { JOB_EXPORT_FILES } from "../../common/constants/storage-path.constant";
import { EXPORT_MIME_TYPES } from "../../common/constants/export-type.constant";
import {
buildJobDataFilePath,
buildJobCsvZipPath,
ensureJobExportStructure,
} from "../../common/helpers/file.helper";
import { BaseExportService } from "./base-export.service";
import { extractDomain } from "../../common/helpers/url.helper";
import {
extractMainContent,
stripMarkdown,
} from "../../common/helpers/data-contract.helper";
import { extractMainContent } from "../../common/helpers/data-contract.helper";
export class CsvExportService extends BaseExportService {
readonly mimeType = EXPORT_MIME_TYPES.CSV;
......@@ -25,39 +24,57 @@ export class CsvExportService extends BaseExportService {
protected async executeExport(
job: CrawlJob & { pages: CrawlPage[] },
): Promise<{ fileName: string; filePath: string }> {
// Fetch assets once — reused for links.csv and images.csv
const assets = await this.crawlAssetRepository.findByJobId(job.id);
// 1. pages.csv
await this.exportPages(job);
// 2. links.csv
await this.exportLinks(job, assets);
// 3. images.csv
await this.exportImages(job, assets);
// 4. Bundle all CSVs into csv.zip
return this.zipCsvFolder(job.id);
}
async exportPages(
job: CrawlJob & { pages: CrawlPage[] },
): Promise<{ fileName: string; filePath: string }> {
ensureJobExportStructure(job.id);
const { fileName, filePath } = buildJobDataFilePath(
job.id,
JOB_EXPORT_FILES.PAGES_CSV,
);
const headers = [
"pageId",
"url",
"title",
"description",
"status",
"statusCode",
"rawMarkdown",
"cleanText",
"errorMessage",
"wordCount",
"dataQualityScore",
"mainContent",
"crawledAt",
];
const rows = job.pages.map((page) => {
const rawMarkdown = page.markdownContent ?? "";
const mainContent = extractMainContent(rawMarkdown);
const cleanText = mainContent ? stripMarkdown(mainContent) : "";
return [
page.id,
this.escapeCsv(page.url),
this.escapeCsv(page.title ?? ""),
this.escapeCsv(page.description ?? ""),
this.escapeCsv(page.status),
page.statusCode ?? "",
this.escapeCsv(rawMarkdown),
this.escapeCsv(cleanText),
this.escapeCsv(page.errorMessage ?? ""),
page.wordCount ?? 0,
page.dataQualityScore ?? "",
this.escapeCsv(mainContent),
page.crawledAt?.toISOString() ?? "",
];
......@@ -66,11 +83,37 @@ export class CsvExportService extends BaseExportService {
const csv = [headers, ...rows].map((r) => r.join(",")).join("\n");
fs.writeFileSync(filePath, csv, "utf-8");
// 2. links.csv
await this.exportLinks(job, assets);
return { fileName, filePath };
}
// 3. images.csv
await this.exportImages(job, assets);
private async zipCsvFolder(
jobId: string,
): Promise<{ fileName: string; filePath: string }> {
const { fileName, filePath } = buildJobCsvZipPath(jobId);
const filesToZip = [
JOB_EXPORT_FILES.PAGES_CSV,
JOB_EXPORT_FILES.LINKS_CSV,
JOB_EXPORT_FILES.IMAGES_CSV,
];
await new Promise<void>((resolve, reject) => {
const output = fs.createWriteStream(filePath);
const archive = archiver("zip", { zlib: { level: 9 } });
output.on("close", resolve);
archive.on("error", reject);
archive.pipe(output);
for (const file of filesToZip) {
const { filePath: srcPath } = buildJobDataFilePath(jobId, file);
if (fs.existsSync(srcPath)) {
archive.file(srcPath, { name: file });
}
}
archive.finalize();
});
return { fileName, filePath };
}
......
import fs from "fs";
import archiver from "archiver";
import { parse as parseHtml } from "node-html-parser";
import {
CrawlJob,
......@@ -12,6 +13,7 @@ import {
buildJobDataFilePath,
buildJobDataRawFilePath,
buildJobDataCleanFilePath,
buildJobJsonZipPath,
} from "../../common/helpers/file.helper";
import { BaseExportService } from "./base-export.service";
import { extractDomain } from "../../common/helpers/url.helper";
......@@ -116,7 +118,84 @@ export class JsonExportService extends BaseExportService {
"utf-8",
);
return { fileName, filePath };
this.writeStructuredJson(job);
return this.zipJsonFolder(job.id);
}
public writeStructuredJson(job: CrawlJob & { pages: CrawlPage[] }): void {
const { filePath } = buildJobDataFilePath(
job.id,
JOB_EXPORT_FILES.STRUCTURED_JSON,
);
const records = job.pages
.filter((p) => p.structuredData != null)
.map((p) => ({
pageId: p.id,
url: p.url,
structuredData: p.structuredData,
}));
if (records.length > 0) {
fs.writeFileSync(
filePath,
JSON.stringify({ jobId: job.id, records }, null, 2),
"utf-8",
);
}
}
public async zipJsonFolder(
jobId: string,
): Promise<{ fileName: string; filePath: string }> {
const { fileName, filePath: zipPath } = buildJobJsonZipPath(jobId);
const { filePath: pagesJsonPath } = buildJobDataFilePath(
jobId,
JOB_EXPORT_FILES.PAGES_JSON,
);
const { filePath: rawJsonPath } = buildJobDataRawFilePath(
jobId,
JOB_EXPORT_FILES.PAGES_RAW_JSON,
);
const { filePath: cleanJsonPath } = buildJobDataCleanFilePath(
jobId,
JOB_EXPORT_FILES.PAGES_CLEAN_JSON,
);
const { filePath: structuredJsonPath } = buildJobDataFilePath(
jobId,
JOB_EXPORT_FILES.STRUCTURED_JSON,
);
return new Promise((resolve, reject) => {
const output = fs.createWriteStream(zipPath);
const archive = archiver("zip", { zlib: { level: 9 } });
output.on("close", () => resolve({ fileName, filePath: zipPath }));
archive.on("error", (err) => reject(err));
archive.pipe(output);
if (fs.existsSync(pagesJsonPath)) {
archive.file(pagesJsonPath, { name: JOB_EXPORT_FILES.PAGES_JSON });
}
if (fs.existsSync(cleanJsonPath)) {
archive.file(cleanJsonPath, {
name: `clean/${JOB_EXPORT_FILES.PAGES_CLEAN_JSON}`,
});
}
if (fs.existsSync(rawJsonPath)) {
archive.file(rawJsonPath, {
name: `raw/${JOB_EXPORT_FILES.PAGES_RAW_JSON}`,
});
}
if (fs.existsSync(structuredJsonPath)) {
archive.file(structuredJsonPath, {
name: JOB_EXPORT_FILES.STRUCTURED_JSON,
});
}
archive.finalize();
});
}
private parsePageTables(
......
......@@ -4,7 +4,6 @@ import { CrawlJob, CrawlPage } from "../../common/types/database.types";
import { JOB_EXPORT_SUBDIRS } from "../../common/constants/storage-path.constant";
import { EXPORT_MIME_TYPES } from "../../common/constants/export-type.constant";
import {
buildJobMarkdownFilePath,
buildJobMarkdownRawFilePath,
buildJobMarkdownCleanFilePath,
buildJobMarkdownZipPath,
......@@ -36,22 +35,13 @@ export class MarkdownExportService extends BaseExportService {
page.markdownContent ??
`# ${page.title ?? page.url}\n\n**URL:** ${page.url}\n\nNo content available.`;
// 1. Tương thích ngược: ghi ở markdown/
const { fileName, filePath } = buildJobMarkdownFilePath(
job.id,
idx,
page.url,
);
fs.writeFileSync(filePath, rawContent, "utf-8");
results.push({ fileName, filePath });
// 2. Ghi bản raw ở markdown/raw/
// 1. Ghi bản raw ở markdown/raw/
const { fileName: rawName, filePath: rawPath } =
buildJobMarkdownRawFilePath(job.id, idx, page.url);
fs.writeFileSync(rawPath, rawContent, "utf-8");
results.push({ fileName: rawName, filePath: rawPath });
// 3. Ghi bản clean ở markdown/clean/
// 2. Ghi bản clean ở markdown/clean/
const { fileName: cleanName, filePath: cleanPath } =
buildJobMarkdownCleanFilePath(job.id, idx, page.url);
const cleanContent =
......
This diff is collapsed.
......@@ -75,11 +75,14 @@ export class ZipExportService extends BaseExportService {
url: p.url,
structuredData: p.structuredData,
}));
fs.writeFileSync(
filePath,
JSON.stringify({ jobId: job.id, records }, null, 2),
"utf-8",
);
if (records.length > 0) {
fs.writeFileSync(
filePath,
JSON.stringify({ jobId: job.id, records }, null, 2),
"utf-8",
);
}
}
private writeMetadata(job: CrawlJob): void {
......
Markdown is supported
0% or
You are about to add 0 people to the discussion. Proceed with caution.
Finish editing this message first!
Please register or to comment