A hard dollar limit per scrape job with model_instance (ChatOpenAI base_url + default_headers) #1165
domondi1
started this conversation in
Show and tell
Replies: 0 comments
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
A ScrapeGraphAI run can be one call or many:
SmartScraperMultiGraphover a list of URLs, large pages split into chunks, merge steps. If you run scrapes as jobs (per customer, per batch) and want a hard dollar ceiling per job, themodel_instanceoption is enough. Pass a LangChainChatOpenAIpointed at a local gateway, with the job id and budget as default headers. I maintain the gateway (Inferrail, open source).Every LLM call in that job carries the same id, so they share the $1.50. A call that would go past it gets HTTP 402 before it reaches OpenAI. Then
inferrail work catalog-batch-17prints calls, tokens and dollars for the job, andinferrail report --by work_idlists all jobs.What I checked (scrapegraphai 2.3.0, langchain-openai 1.6.7, inferrail 0.4.15,
SmartScraperGraphon raw HTML against a stand-in OpenAI endpoint with fixed token counts):X-Inferrail-*headers are not forwarded upstream.graph.run()raisedopenai.APIStatusError402 (INFERRAIL_E010, with the budget and the estimate in the message) after one attempt, and nothing reached the upstream. Catch it if you want partial results instead.Caveats: a dollar budget needs a price for the model (most OpenAI and Anthropic models are built in,
inferrail models; an unpriced model is refused, not guessed), and the gateway stores usage and cost per call, never page content or prompts. I only tested the single-page graph; the multi graph uses the samemodel_instance.All reactions