跳转到内容
搜索文档

使用 AI 分析数据

最后更新 查看 MarkdownAgent 设置

构建一个 AI 驱动的数据分析系统:接受 CSV 上传,使用 Claude 生成 Python 分析代码,在沙箱中执行,并返回可视化结果。

预计完成时间:25 分钟

前提条件

  1. 注册 Cloudflare 账户
  2. 安装 Node.js

Node.js 版本管理器

使用 Voltanvm 等 Node 版本管理器,以避免权限问题并切换 Node.js 版本。本指南后续将介绍的 Wrangler 需要 Node 版本 16.17.0 或更高。

你还需要:

1. 创建项目

创建新的 Sandbox SDK 项目:

npm create cloudflare@latest -- analyze-data --template=cloudflare/sandbox-sdk/examples/minimal
cd analyze-data

2. 安装依赖

npm i @anthropic-ai/sdk

3. 构建分析处理程序

替换 src/index.ts

import { getSandbox, proxyToSandbox, type Sandbox } from "@cloudflare/sandbox";
import Anthropic from "@anthropic-ai/sdk";

export { Sandbox } from "@cloudflare/sandbox";

interface Env {
	Sandbox: DurableObjectNamespace<Sandbox>;
	ANTHROPIC_API_KEY: string;
}

export default {
	async fetch(request: Request, env: Env): Promise<Response> {
		const proxyResponse = await proxyToSandbox(request, env);
		if (proxyResponse) return proxyResponse;

		if (request.method !== "POST") {
			return Response.json(
				{ error: "POST CSV file and question" },
				{ status: 405 },
			);
		}

		try {
			const formData = await request.formData();
			const csvFile = formData.get("file") as File;
			const question = formData.get("question") as string;

			if (!csvFile || !question) {
				return Response.json(
					{ error: "Missing file or question" },
					{ status: 400 },
				);
			}

			// Upload CSV to sandbox
			const sandbox = getSandbox(env.Sandbox, `analysis-${Date.now()}`);
			const csvPath = "/workspace/data.csv";
			await sandbox.writeFile(csvPath, await csvFile.text());

			// Analyze CSV structure
			const structure = await sandbox.exec(
				`python3 -c "import pandas as pd; df = pd.read_csv('${csvPath}'); print(f'Rows: {len(df)}'); print(f'Columns: {list(df.columns)[:5]}')"`,
			);

			if (!structure.success) {
				return Response.json(
					{ error: "Failed to read CSV", details: structure.stderr },
					{ status: 400 },
				);
			}

			// Generate analysis code with Claude
			const code = await generateAnalysisCode(
				env.ANTHROPIC_API_KEY,
				csvPath,
				question,
				structure.stdout,
			);

			// Write and execute the analysis code
			await sandbox.writeFile("/workspace/analyze.py", code);
			const result = await sandbox.exec("python /workspace/analyze.py");

			if (!result.success) {
				return Response.json(
					{ error: "Analysis failed", details: result.stderr },
					{ status: 500 },
				);
			}

			async function streamToBase64(stream) {
			  const blob = await new Response(stream).blob();
			  const buffer = await blob.arrayBuffer();
			  const bytes = new Uint8Array(buffer);

			  // Convert to base64
			  let binary = '';
			  for (let i = 0; i < bytes.length; i++) {
			    binary += String.fromCharCode(bytes[i]);
			  }
			  return btoa(binary);
			}

			// Check for generated chart
			let chart = null;
			try {
				const { content, mimeType } = await sandbox.readFile("/workspace/chart.png", {
					encoding: "none"
				});
				chart = `data:${mimeType};base64,${await streamToBase64(content)}`;
			} catch {
				// No chart generated
			}

			await sandbox.destroy();

			return Response.json({
				success: true,
				output: result.stdout,
				chart,
				code,
			});
		} catch (error: any) {
			return Response.json({ error: error.message }, { status: 500 });
		}
	},
};

async function generateAnalysisCode(
	apiKey: string,
	csvPath: string,
	question: string,
	csvStructure: string,
): Promise<string> {
	const anthropic = new Anthropic({ apiKey });

	const response = await anthropic.messages.create({
		model: "claude-sonnet-4-5",
		max_tokens: 2048,
		messages: [
			{
				role: "user",
				content: `CSV at ${csvPath}:
${csvStructure}

Question: "${question}"

Generate Python code that:
- Reads CSV with pandas
- Answers the question
- Saves charts to /workspace/chart.png if helpful
- Prints findings to stdout

Use pandas, numpy, matplotlib.`,
			},
		],
		tools: [
			{
				name: "generate_python_code",
				description: "Generate Python code for data analysis",
				input_schema: {
					type: "object",
					properties: {
						code: { type: "string", description: "Complete Python code" },
					},
					required: ["code"],
				},
			},
		],
	});

	for (const block of response.content) {
		if (block.type === "tool_use" && block.name === "generate_python_code") {
			return (block.input as { code: string }).code;
		}
	}

	throw new Error("Failed to generate code");
}

4. 设置本地环境变量

在项目根目录创建 .dev.vars 文件,用于本地开发:

echo "ANTHROPIC_API_KEY=your_api_key_here\nSANDBOX_TRANSPORT=rpc" > .dev.vars

your_api_key_here 替换为你在 Anthropic Console 中的实际 API key。

SANDBOX_TRANSPORT 是使用新文件流式 API 所必需的。

5. 本地测试

下载示例 CSV:

# Create a test CSV
echo "year,rating,title
2020,8.5,Movie A
2021,7.2,Movie B
2022,9.1,Movie C" > test.csv

启动开发服务器:

npm run dev

使用 curl 测试:

curl -X POST http://localhost:8787 \
  -F "[email protected]" \
  -F "question=What is the average rating by year?"

响应:

{
	"success": true,
	"output": "Average ratings by year:\n2020: 8.5\n2021: 7.2\n2022: 9.1",
	"chart": "data:image/png;base64,...",
	"code": "import pandas as pd\nimport matplotlib.pyplot as plt\n..."
}

6. 部署

部署你的 Worker:

npx wrangler deploy

然后将 Anthropic API key 设为生产环境 secret:

npx wrangler secret put ANTHROPIC_API_KEY

出现提示时,粘贴来自 Anthropic Console 的 API key。

你构建了什么

一个 AI 数据分析系统,它能够:

  • 将 CSV 文件上传到沙箱
  • 使用 Claude 的工具调用生成分析代码
  • 使用 pandas 和 matplotlib 执行 Python
  • 返回文本输出和可视化结果

后续步骤

这篇文档对您有帮助吗?