---------------------------------------------------------------------- This is the API documentation for the oxyscraper library. ---------------------------------------------------------------------- ## Classes Core classes Amazon(*, source: Literal['amazon'] = 'amazon', url: str, render: Optional[Literal['html', 'png', '']] = None, user_agent_type: Optional[Literal['desktop', 'mobile', 'mobile_android', 'mobile_ios', 'tablet', 'tablet_android', 'tablet_ios', 'desktop_chrome', 'desktop_edge', 'desktop_firefox', 'desktop_opera', 'desktop_safari']] = None, callback_url: str | None = None, parse: bool | None = None, markdown: bool | None = None, xhr: bool | None = None, parser_preset: str | None = None, content_encoding: Optional[Literal['base64', 'utf-8']] = None, client_notes: str | None = None, aggregate_name: str | None = None, geo_location: str | None = None, locale: Optional[Literal['ar_AE', 'bn_IN', 'cs_CZ', 'da_DK', 'de_DE', 'de_US', 'en_AE', 'en_AU', 'en_CA', 'en_GB', 'en_IE', 'en_IN', 'en_SG', 'en_US', 'en_ZA', 'es_ES', 'es_MX', 'es_US', 'fr_BE', 'fr_CA', 'fr_FR', 'he_IL', 'hi_IN', 'it_IT', 'ja_JP', 'kn_IN', 'ko_KR', 'ml_IN', 'mr_IN', 'nl_BE', 'nl_NL', 'pl_PL', 'pt_BR', 'pt_PT', 'sv_SE', 'ta_IN', 'te_IN', 'tr_TR', 'zh_CN', 'zh_TW']] = None, context: list[oxyscraper._payloads._ContextItem] | None = None, storage_type: Optional[Literal['gcs', 's3', 'tos', 's3_compatible']] = None, storage_url: str | None = None, parsing_instructions: Optional[Annotated[ParsingInstructions, PlainValidator(func=, json_schema_input_type=Any), WithJsonSchema(json_schema={}, mode=None)]] = None, browser_instructions: list[typing.Annotated[oxyscraper._payloads._Click | oxyscraper._payloads._Input | oxyscraper._payloads._Scroll | oxyscraper._payloads._Wait | oxyscraper._payloads._FetchResource, FieldInfo(annotation=NoneType, required=True, discriminator='type')]] | None = None, extra: dict[str, typing.Any] = {}, currency: Optional[Literal['AED', 'AMD', 'ARS', 'AUD', 'AWG', 'AZN', 'BBD', 'BGN', 'BHD', 'BMD', 'BND', 'BOB', 'BRL', 'BSD', 'BZD', 'CAD', 'CHF', 'CLP', 'CNY', 'COP', 'CRC', 'CZK', 'DKK', 'DOP', 'EGP', 'EUR', 'GBP', 'GHS', 'GTQ', 'HKD', 'HNL', 'HUF', 'IDR', 'ILS', 'INR', 'JMD', 'JOD', 'JPY', 'KES', 'KHR', 'KRW', 'KWD', 'KYD', 'KZT', 'LBP', 'LKR', 'MAD', 'MNT', 'MOP', 'MUR', 'MXN', 'MYR', 'NAD', 'NGN', 'NOK', 'NZD', 'OMR', 'PAB', 'PEN', 'PHP', 'PKR', 'PLN', 'PYG', 'QAR', 'RON', 'RUB', 'SAR', 'SEK', 'SGD', 'THB', 'TRY', 'TTD', 'TWD', 'TZS', 'USD', 'UYU', 'VND', 'XCD', 'ZAR']] = None) -> None An `amazon` job, which scrapes an Amazon URL. The API runs a product URL as an `amazon_product` job and a search URL as an `amazon_search` job, and appends `language=` to every URL. A Best Sellers URL may render and bill as a rendered result. The model takes no `domain`, because the URL's host sets it, and no `start_page` or `pages`, because the API bills them with no effect. [What a live test shows about the Amazon models](https://github.com/ozanozbeker/oxyscraper/blob/main/docs/research/live-amazon.md) records the run that checked it. Attributes ---------- url An Amazon URL. The API rejects a URL outside Amazon for free, and rejects `parse=True` for free on a page type without a dedicated parser. locale The page's language, which the API rejects for free unless the host lists it. Without it, `amazon.ae` returns its Arabic page. geo_location A postal code inside the host's country, or an ISO 3166-1 alpha-2 code outside it, which the API rejects for free if it does not fit. `ae`, `com.be`, `eg`, `ie`, `pl`, `sa`, `se` and `sg` run no check, so a wrong value may bill. `99999` on `com` faults after 120 seconds. currency The currency of the prices, which the API rejects for free unless the host lists it. On `amazon.com`, any currency but USD raises without a 2-letter country code in `geo_location`, because the API bills prices in USD then. user_agent_type It raises with `render`, because a rendered job ignores it and still bills. AmazonBestsellers(*, source: Literal['amazon_bestsellers'] = 'amazon_bestsellers', query: typing.Annotated[str, _PydanticGeneralMetadata(pattern='^[0-9]+$')], render: Optional[Literal['html', 'png', '']] = None, callback_url: str | None = None, parse: bool | None = None, start_page: Optional[Annotated[int, Gt(gt=0)]] = None, pages: Optional[Annotated[int, Gt(gt=0)]] = None, markdown: bool | None = None, xhr: bool | None = None, parser_preset: str | None = None, content_encoding: Optional[Literal['base64', 'utf-8']] = None, client_notes: str | None = None, aggregate_name: str | None = None, geo_location: str | None = None, locale: Optional[Literal['ar_AE', 'bn_IN', 'cs_CZ', 'da_DK', 'de_DE', 'de_US', 'en_AE', 'en_AU', 'en_CA', 'en_GB', 'en_IE', 'en_IN', 'en_SG', 'en_US', 'en_ZA', 'es_ES', 'es_MX', 'es_US', 'fr_BE', 'fr_CA', 'fr_FR', 'he_IL', 'hi_IN', 'it_IT', 'ja_JP', 'kn_IN', 'ko_KR', 'ml_IN', 'mr_IN', 'nl_BE', 'nl_NL', 'pl_PL', 'pt_BR', 'pt_PT', 'sv_SE', 'ta_IN', 'te_IN', 'tr_TR', 'zh_CN', 'zh_TW']] = None, domain: Optional[Literal['ae', 'ca', 'cn', 'co.jp', 'co.uk', 'com', 'com.au', 'com.be', 'com.br', 'com.mx', 'com.tr', 'de', 'eg', 'es', 'fr', 'ie', 'in', 'it', 'nl', 'pl', 'sa', 'se', 'sg', 'co.za']] = None, context: list[oxyscraper._payloads._ContextItem] | None = None, storage_type: Optional[Literal['gcs', 's3', 'tos', 's3_compatible']] = None, storage_url: str | None = None, parsing_instructions: Optional[Annotated[ParsingInstructions, PlainValidator(func=, json_schema_input_type=Any), WithJsonSchema(json_schema={}, mode=None)]] = None, browser_instructions: list[typing.Annotated[oxyscraper._payloads._Click | oxyscraper._payloads._Input | oxyscraper._payloads._Scroll | oxyscraper._payloads._Wait | oxyscraper._payloads._FetchResource, FieldInfo(annotation=NoneType, required=True, discriminator='type')]] | None = None, extra: dict[str, typing.Any] = {}, currency: Optional[Literal['AED', 'AMD', 'ARS', 'AUD', 'AWG', 'AZN', 'BBD', 'BGN', 'BHD', 'BMD', 'BND', 'BOB', 'BRL', 'BSD', 'BZD', 'CAD', 'CHF', 'CLP', 'CNY', 'COP', 'CRC', 'CZK', 'DKK', 'DOP', 'EGP', 'EUR', 'GBP', 'GHS', 'GTQ', 'HKD', 'HNL', 'HUF', 'IDR', 'ILS', 'INR', 'JMD', 'JOD', 'JPY', 'KES', 'KHR', 'KRW', 'KWD', 'KYD', 'KZT', 'LBP', 'LKR', 'MAD', 'MNT', 'MOP', 'MUR', 'MXN', 'MYR', 'NAD', 'NGN', 'NOK', 'NZD', 'OMR', 'PAB', 'PEN', 'PHP', 'PKR', 'PLN', 'PYG', 'QAR', 'RON', 'RUB', 'SAR', 'SEK', 'SGD', 'THB', 'TRY', 'TTD', 'TWD', 'TZS', 'USD', 'UYU', 'VND', 'XCD', 'ZAR']] = None) -> None An `amazon_bestsellers` job, which scrapes the Best Sellers page of one browse node. The API renders every job, and bills each result as rendered, although it takes nothing from the rendered limit. `render=""` turns forced rendering off, and the job then faults. The model takes no `user_agent_type`, because rendering ignores it. [What a live test shows about the Amazon models](https://github.com/ozanozbeker/oxyscraper/blob/main/docs/research/live-amazon.md) records the run that checked it. Attributes ---------- query A browse node ID, such as `172541`. Any other value raises, because the API bills a rendered "Best undefined" page for it. An unknown node ID bills that page too. domain The marketplace, `com` by default. `co.za` works, although no docs page lists it, and every `cn` job faulted. locale The page's language, which the API rejects for free unless the domain lists it. Without it, `ae` returns its Arabic page. geo_location A postal code inside the domain's country, or an ISO 3166-1 alpha-2 code outside it, which the API rejects for free if it does not fit. `ae`, `com.be`, `eg`, `ie`, `pl`, `sa`, `se` and `sg` run no check, so a wrong value may bill. `99999` on `com` faults after 120 seconds. start_page A page past the last faults. pages The API rejects more than 20 for free. currency The currency of the prices, which the API rejects for free unless the domain lists it. On `com`, any currency but USD raises without a 2-letter country code in `geo_location`, because the API bills prices in USD then. AmazonPricing(*, source: Literal['amazon_pricing'] = 'amazon_pricing', query: typing.Annotated[str, MaxLen(max_length=10)], render: Optional[Literal['html', 'png', '']] = None, user_agent_type: Optional[Literal['desktop', 'mobile', 'mobile_android', 'mobile_ios', 'tablet', 'tablet_android', 'tablet_ios', 'desktop_chrome', 'desktop_edge', 'desktop_firefox', 'desktop_opera', 'desktop_safari']] = None, callback_url: str | None = None, parse: bool | None = None, start_page: Optional[Annotated[int, Gt(gt=0)]] = None, pages: Optional[Annotated[int, Gt(gt=0)]] = None, markdown: bool | None = None, xhr: bool | None = None, parser_preset: str | None = None, content_encoding: Optional[Literal['base64', 'utf-8']] = None, client_notes: str | None = None, aggregate_name: str | None = None, geo_location: str | None = None, locale: Optional[Literal['ar_AE', 'bn_IN', 'cs_CZ', 'da_DK', 'de_DE', 'de_US', 'en_AE', 'en_AU', 'en_CA', 'en_GB', 'en_IE', 'en_IN', 'en_SG', 'en_US', 'en_ZA', 'es_ES', 'es_MX', 'es_US', 'fr_BE', 'fr_CA', 'fr_FR', 'he_IL', 'hi_IN', 'it_IT', 'ja_JP', 'kn_IN', 'ko_KR', 'ml_IN', 'mr_IN', 'nl_BE', 'nl_NL', 'pl_PL', 'pt_BR', 'pt_PT', 'sv_SE', 'ta_IN', 'te_IN', 'tr_TR', 'zh_CN', 'zh_TW']] = None, domain: Optional[Literal['ae', 'ca', 'cn', 'co.jp', 'co.uk', 'com', 'com.au', 'com.be', 'com.br', 'com.mx', 'com.tr', 'de', 'eg', 'es', 'fr', 'ie', 'in', 'it', 'nl', 'pl', 'sa', 'se', 'sg', 'co.za']] = None, context: list[oxyscraper._payloads._ContextItem] | None = None, storage_type: Optional[Literal['gcs', 's3', 'tos', 's3_compatible']] = None, storage_url: str | None = None, parsing_instructions: Optional[Annotated[ParsingInstructions, PlainValidator(func=, json_schema_input_type=Any), WithJsonSchema(json_schema={}, mode=None)]] = None, browser_instructions: list[typing.Annotated[oxyscraper._payloads._Click | oxyscraper._payloads._Input | oxyscraper._payloads._Scroll | oxyscraper._payloads._Wait | oxyscraper._payloads._FetchResource, FieldInfo(annotation=NoneType, required=True, discriminator='type')]] | None = None, extra: dict[str, typing.Any] = {}) -> None An `amazon_pricing` job, which scrapes the offers of one product. The model takes no `currency`, because the API bills it with no effect. [What a live test shows about the Amazon models](https://github.com/ozanozbeker/oxyscraper/blob/main/docs/research/live-amazon.md) records the run that checked it. Attributes ---------- query An ASIN. One longer than 10 characters raises, because the API bills a 404 page for it. The API rejects a shorter one, or one with lowercase letters, for free, and bills a 404 page for one that does not exist. domain The marketplace, `com` by default. `co.za` works, although no docs page lists it, and every `cn` job faulted. locale The page's language, which the API rejects for free unless the domain lists it. Without it, `ae` returns its Arabic page. geo_location A postal code inside the domain's country, or an ISO 3166-1 alpha-2 code outside it, which the API rejects for free if it does not fit. `ae`, `com.be`, `eg`, `ie`, `pl`, `sa`, `se` and `sg` run no check, so a wrong value may bill. `99999` on `com` faults after 120 seconds. start_page A page past the last bills. pages The API rejects more than 20 for free. user_agent_type It raises with `render`, because a rendered job ignores it and still bills. AmazonProduct(*, source: Literal['amazon_product'] = 'amazon_product', query: typing.Annotated[str, MaxLen(max_length=10)], render: Optional[Literal['html', 'png', '']] = None, user_agent_type: Optional[Literal['desktop', 'mobile', 'mobile_android', 'mobile_ios', 'tablet', 'tablet_android', 'tablet_ios', 'desktop_chrome', 'desktop_edge', 'desktop_firefox', 'desktop_opera', 'desktop_safari']] = None, callback_url: str | None = None, parse: bool | None = None, markdown: bool | None = None, xhr: bool | None = None, parser_preset: str | None = None, content_encoding: Optional[Literal['base64', 'utf-8']] = None, client_notes: str | None = None, aggregate_name: str | None = None, geo_location: str | None = None, locale: Optional[Literal['ar_AE', 'bn_IN', 'cs_CZ', 'da_DK', 'de_DE', 'de_US', 'en_AE', 'en_AU', 'en_CA', 'en_GB', 'en_IE', 'en_IN', 'en_SG', 'en_US', 'en_ZA', 'es_ES', 'es_MX', 'es_US', 'fr_BE', 'fr_CA', 'fr_FR', 'he_IL', 'hi_IN', 'it_IT', 'ja_JP', 'kn_IN', 'ko_KR', 'ml_IN', 'mr_IN', 'nl_BE', 'nl_NL', 'pl_PL', 'pt_BR', 'pt_PT', 'sv_SE', 'ta_IN', 'te_IN', 'tr_TR', 'zh_CN', 'zh_TW']] = None, domain: Optional[Literal['ae', 'ca', 'cn', 'co.jp', 'co.uk', 'com', 'com.au', 'com.be', 'com.br', 'com.mx', 'com.tr', 'de', 'eg', 'es', 'fr', 'ie', 'in', 'it', 'nl', 'pl', 'sa', 'se', 'sg', 'co.za']] = None, context: list[oxyscraper._payloads._ContextItem] | None = None, storage_type: Optional[Literal['gcs', 's3', 'tos', 's3_compatible']] = None, storage_url: str | None = None, parsing_instructions: Optional[Annotated[ParsingInstructions, PlainValidator(func=, json_schema_input_type=Any), WithJsonSchema(json_schema={}, mode=None)]] = None, browser_instructions: list[typing.Annotated[oxyscraper._payloads._Click | oxyscraper._payloads._Input | oxyscraper._payloads._Scroll | oxyscraper._payloads._Wait | oxyscraper._payloads._FetchResource, FieldInfo(annotation=NoneType, required=True, discriminator='type')]] | None = None, extra: dict[str, typing.Any] = {}, autoselect_variant: bool | None = None, currency: Optional[Literal['AED', 'AMD', 'ARS', 'AUD', 'AWG', 'AZN', 'BBD', 'BGN', 'BHD', 'BMD', 'BND', 'BOB', 'BRL', 'BSD', 'BZD', 'CAD', 'CHF', 'CLP', 'CNY', 'COP', 'CRC', 'CZK', 'DKK', 'DOP', 'EGP', 'EUR', 'GBP', 'GHS', 'GTQ', 'HKD', 'HNL', 'HUF', 'IDR', 'ILS', 'INR', 'JMD', 'JOD', 'JPY', 'KES', 'KHR', 'KRW', 'KWD', 'KYD', 'KZT', 'LBP', 'LKR', 'MAD', 'MNT', 'MOP', 'MUR', 'MXN', 'MYR', 'NAD', 'NGN', 'NOK', 'NZD', 'OMR', 'PAB', 'PEN', 'PHP', 'PKR', 'PLN', 'PYG', 'QAR', 'RON', 'RUB', 'SAR', 'SEK', 'SGD', 'THB', 'TRY', 'TTD', 'TWD', 'TZS', 'USD', 'UYU', 'VND', 'XCD', 'ZAR']] = None) -> None An `amazon_product` job, which scrapes one product page. The model takes no `start_page` or `pages`, because the API bills them with no effect. [What a live test shows about the Amazon models](https://github.com/ozanozbeker/oxyscraper/blob/main/docs/research/live-amazon.md) records the run that checked it. Attributes ---------- query An ASIN. One longer than 10 characters raises, because the API bills a 404 page for it. The API rejects a shorter one, or one with lowercase letters, for free, and bills a 404 page for one that does not exist. domain The marketplace, `com` by default. `co.za` works, although no docs page lists it, and every `cn` job faulted. locale The page's language, which the API rejects for free unless the domain lists it. Without it, `ae` returns its Arabic page. geo_location A postal code inside the domain's country, or an ISO 3166-1 alpha-2 code outside it, which the API rejects for free if it does not fit. `ae`, `com.be`, `eg`, `ie`, `pl`, `sa`, `se` and `sg` run no check, so a wrong value may bill. `99999` on `com` faults after 120 seconds. user_agent_type It raises with `render`, because a rendered job ignores it and still bills. autoselect_variant Adds `th=1&psc=1` to the product URL, so the page shows a variant's price and buybox. currency The currency of the prices, which the API rejects for free unless the domain lists it. On `com`, any currency but USD raises without a 2-letter country code in `geo_location`, because the API bills prices in USD then. AmazonSearch(*, source: Literal['amazon_search'] = 'amazon_search', query: str, category_id: str | None = None, render: Optional[Literal['html', 'png', '']] = None, user_agent_type: Optional[Literal['desktop', 'mobile', 'mobile_android', 'mobile_ios', 'tablet', 'tablet_android', 'tablet_ios', 'desktop_chrome', 'desktop_edge', 'desktop_firefox', 'desktop_opera', 'desktop_safari']] = None, callback_url: str | None = None, parse: bool | None = None, start_page: Optional[Annotated[int, Gt(gt=0)]] = None, pages: Optional[Annotated[int, Gt(gt=0)]] = None, markdown: bool | None = None, xhr: bool | None = None, parser_preset: str | None = None, content_encoding: Optional[Literal['base64', 'utf-8']] = None, client_notes: str | None = None, aggregate_name: str | None = None, geo_location: str | None = None, locale: Optional[Literal['ar_AE', 'bn_IN', 'cs_CZ', 'da_DK', 'de_DE', 'de_US', 'en_AE', 'en_AU', 'en_CA', 'en_GB', 'en_IE', 'en_IN', 'en_SG', 'en_US', 'en_ZA', 'es_ES', 'es_MX', 'es_US', 'fr_BE', 'fr_CA', 'fr_FR', 'he_IL', 'hi_IN', 'it_IT', 'ja_JP', 'kn_IN', 'ko_KR', 'ml_IN', 'mr_IN', 'nl_BE', 'nl_NL', 'pl_PL', 'pt_BR', 'pt_PT', 'sv_SE', 'ta_IN', 'te_IN', 'tr_TR', 'zh_CN', 'zh_TW']] = None, domain: Optional[Literal['ae', 'ca', 'cn', 'co.jp', 'co.uk', 'com', 'com.au', 'com.be', 'com.br', 'com.mx', 'com.tr', 'de', 'eg', 'es', 'fr', 'ie', 'in', 'it', 'nl', 'pl', 'sa', 'se', 'sg', 'co.za']] = None, context: list[oxyscraper._payloads._ContextItem] | None = None, storage_type: Optional[Literal['gcs', 's3', 'tos', 's3_compatible']] = None, storage_url: str | None = None, parsing_instructions: Optional[Annotated[ParsingInstructions, PlainValidator(func=, json_schema_input_type=Any), WithJsonSchema(json_schema={}, mode=None)]] = None, browser_instructions: list[typing.Annotated[oxyscraper._payloads._Click | oxyscraper._payloads._Input | oxyscraper._payloads._Scroll | oxyscraper._payloads._Wait | oxyscraper._payloads._FetchResource, FieldInfo(annotation=NoneType, required=True, discriminator='type')]] | None = None, extra: dict[str, typing.Any] = {}, currency: Optional[Literal['AED', 'AMD', 'ARS', 'AUD', 'AWG', 'AZN', 'BBD', 'BGN', 'BHD', 'BMD', 'BND', 'BOB', 'BRL', 'BSD', 'BZD', 'CAD', 'CHF', 'CLP', 'CNY', 'COP', 'CRC', 'CZK', 'DKK', 'DOP', 'EGP', 'EUR', 'GBP', 'GHS', 'GTQ', 'HKD', 'HNL', 'HUF', 'IDR', 'ILS', 'INR', 'JMD', 'JOD', 'JPY', 'KES', 'KHR', 'KRW', 'KWD', 'KYD', 'KZT', 'LBP', 'LKR', 'MAD', 'MNT', 'MOP', 'MUR', 'MXN', 'MYR', 'NAD', 'NGN', 'NOK', 'NZD', 'OMR', 'PAB', 'PEN', 'PHP', 'PKR', 'PLN', 'PYG', 'QAR', 'RON', 'RUB', 'SAR', 'SEK', 'SGD', 'THB', 'TRY', 'TTD', 'TWD', 'TZS', 'USD', 'UYU', 'VND', 'XCD', 'ZAR']] = None, sort_by: Optional[Literal['most_recent', 'price_low_to_high', 'price_high_to_low', 'featured', 'average_review', 'bestsellers']] = None, refinements: list[str] | None = None, min_price: Optional[Annotated[int, Gt(gt=0)]] = None, max_price: Optional[Annotated[int, Gt(gt=0)]] = None, merchant_id: str | None = None) -> None An `amazon_search` job, which scrapes the results of one search. [What a live test shows about the Amazon models](https://github.com/ozanozbeker/oxyscraper/blob/main/docs/research/live-amazon.md) records the run that checked it. Attributes ---------- query The search term. domain The marketplace, `com` by default. `co.za` works, although no docs page lists it, and every `cn` job faulted. locale The page's language, which the API rejects for free unless the domain lists it. Without it, `ae` returns its Arabic page. geo_location A postal code inside the domain's country, or an ISO 3166-1 alpha-2 code outside it, which the API rejects for free if it does not fit. `ae`, `com.be`, `eg`, `ie`, `pl`, `sa`, `se` and `sg` run no check, so a wrong value may bill. `99999` on `com` faults after 120 seconds. start_page A page past the last bills. pages The API rejects more than 20 for free. user_agent_type It raises with `render`, because a rendered job ignores it and still bills. currency The currency of the prices, which the API rejects for free unless the domain lists it. On `com`, any currency but USD raises without a 2-letter country code in `geo_location`, because the API bills prices in USD then. sort_by The order of the results. refinements Amazon refinement codes, such as `p_123:256097`. min_price The lowest price, in cents, so `5000` means 50.00. It raises for 0, because the API then applies no filter and still bills. The API rejects a `min_price` above `max_price` for free. max_price The highest price, in cents. category_id A browse node ID that limits the search, sent as a `context` item rather than as the input key of `target_category`. merchant_id A seller ID that limits the search. AmazonSellers(*, source: Literal['amazon_sellers'] = 'amazon_sellers', query: str, render: Optional[Literal['html', 'png', '']] = None, user_agent_type: Optional[Literal['desktop', 'mobile', 'mobile_android', 'mobile_ios', 'tablet', 'tablet_android', 'tablet_ios', 'desktop_chrome', 'desktop_edge', 'desktop_firefox', 'desktop_opera', 'desktop_safari']] = None, callback_url: str | None = None, parse: bool | None = None, markdown: bool | None = None, xhr: bool | None = None, parser_preset: str | None = None, content_encoding: Optional[Literal['base64', 'utf-8']] = None, client_notes: str | None = None, aggregate_name: str | None = None, geo_location: str | None = None, locale: Optional[Literal['ar_AE', 'bn_IN', 'cs_CZ', 'da_DK', 'de_DE', 'de_US', 'en_AE', 'en_AU', 'en_CA', 'en_GB', 'en_IE', 'en_IN', 'en_SG', 'en_US', 'en_ZA', 'es_ES', 'es_MX', 'es_US', 'fr_BE', 'fr_CA', 'fr_FR', 'he_IL', 'hi_IN', 'it_IT', 'ja_JP', 'kn_IN', 'ko_KR', 'ml_IN', 'mr_IN', 'nl_BE', 'nl_NL', 'pl_PL', 'pt_BR', 'pt_PT', 'sv_SE', 'ta_IN', 'te_IN', 'tr_TR', 'zh_CN', 'zh_TW']] = None, domain: Optional[Literal['ae', 'ca', 'cn', 'co.jp', 'co.uk', 'com', 'com.au', 'com.be', 'com.br', 'com.mx', 'com.tr', 'de', 'eg', 'es', 'fr', 'ie', 'in', 'it', 'nl', 'pl', 'sa', 'se', 'sg', 'co.za']] = None, context: list[oxyscraper._payloads._ContextItem] | None = None, storage_type: Optional[Literal['gcs', 's3', 'tos', 's3_compatible']] = None, storage_url: str | None = None, parsing_instructions: Optional[Annotated[ParsingInstructions, PlainValidator(func=, json_schema_input_type=Any), WithJsonSchema(json_schema={}, mode=None)]] = None, browser_instructions: list[typing.Annotated[oxyscraper._payloads._Click | oxyscraper._payloads._Input | oxyscraper._payloads._Scroll | oxyscraper._payloads._Wait | oxyscraper._payloads._FetchResource, FieldInfo(annotation=NoneType, required=True, discriminator='type')]] | None = None, extra: dict[str, typing.Any] = {}) -> None An `amazon_sellers` job, which scrapes one seller's page. The model takes no `start_page` or `pages`, because the API bills them with no effect. [What a live test shows about the Amazon models](https://github.com/ozanozbeker/oxyscraper/blob/main/docs/research/live-amazon.md) records the run that checked it. Attributes ---------- query A seller ID, such as `A2OL0VKAHK1LYK`. The API bills a 404 page for one that does not exist. domain The marketplace, `com` by default. `co.za` works, although no docs page lists it, and every `cn` job faulted. locale The page's language, which the API rejects for free unless the domain lists it. Without it, `ae` returns its Arabic page. geo_location A postal code inside the domain's country, or an ISO 3166-1 alpha-2 code outside it, which the API rejects for free if it does not fit. `ae`, `com.be`, `eg`, `ie`, `pl`, `sa`, `se` and `sg` run no check, so a wrong value may bill. `99999` on `com` faults after 120 seconds. user_agent_type It raises with `render`, because a rendered job ignores it and still bills. Payload(*, source: str, query: str | None = None, url: str | None = None, product_id: str | None = None, prompt: str | None = None, video_id: str | None = None, channel_handle: str | None = None, category_id: str | None = None, render: Optional[Literal['html', 'png', '']] = None, user_agent_type: Optional[Literal['desktop', 'mobile', 'mobile_android', 'mobile_ios', 'tablet', 'tablet_android', 'tablet_ios', 'desktop_chrome', 'desktop_edge', 'desktop_firefox', 'desktop_opera', 'desktop_safari']] = None, callback_url: str | None = None, parse: bool | None = None, start_page: Optional[Annotated[int, Gt(gt=0)]] = None, pages: Optional[Annotated[int, Gt(gt=0)]] = None, limit: int | None = None, markdown: bool | None = None, xhr: bool | None = None, parser_preset: str | None = None, content_encoding: Optional[Literal['base64', 'utf-8']] = None, client_notes: str | None = None, aggregate_name: str | None = None, geo_location: str | None = None, locale: str | None = None, domain: str | None = None, context: list[oxyscraper._payloads._ContextItem] | None = None, storage_type: Optional[Literal['gcs', 's3', 'tos', 's3_compatible']] = None, storage_url: str | None = None, parsing_instructions: Optional[Annotated[ParsingInstructions, PlainValidator(func=, json_schema_input_type=Any), WithJsonSchema(json_schema={}, mode=None)]] = None, browser_instructions: list[typing.Annotated[oxyscraper._payloads._Click | oxyscraper._payloads._Input | oxyscraper._payloads._Scroll | oxyscraper._payloads._Wait | oxyscraper._payloads._FetchResource, FieldInfo(annotation=NoneType, required=True, discriminator='type')]] | None = None, extra: dict[str, typing.Any] = {}, **extra_data: Any) -> None One job's body, for any source. `Payload` types each parameter that keeps one name, placement, type and value set on every source that takes it. Any other keyword goes into the body as it is, so a source without a model of its own still runs. An unset field stays out of the body, so the API applies its own default. A payload sets exactly one input key, to a non-empty string. Otherwise it raises only for a mistake that the API would bill, and leaves each free check to the API. Attributes ---------- source The source that runs the job. It takes any string, because the API rejects an unknown source for free. query The input of most search and product sources. url The input of `universal` and of the sources that take a page's URL. product_id The input of product sources such as `walmart_product`. prompt The input of `chatgpt`, `gemini` and `perplexity`. video_id The input of `youtube_video_trainability`. channel_handle The input of `youtube_channel`. category_id The input of `target_category`. render `html` or `png` renders the page in a browser, and `""` turns off forced rendering. user_agent_type The device of the job's user agent. A `desktop_*` value draws from the same agents as `desktop`. callback_url The URL that the API calls when the job finishes. parse Returns parsed content, which needs a dedicated parser, `parsing_instructions` or `parser_preset`. start_page The first page to fetch. pages The number of pages to fetch, each billed as one result. limit The number of results on each page, or of videos on `youtube_channel`. markdown Makes Markdown the default output type. xhr Makes the page's Fetch and XHR requests the default output type, and needs `render`. parser_preset The parser preset to parse with, which needs `parse`. content_encoding `base64` returns an image as Base64 text. client_notes Text that the API saves with the job. aggregate_name The Result Aggregator that receives the result. geo_location The location that the job appears to come from, in a format that depends on the source. locale The language of the page, such as `en_US` on Amazon or `de-DE` on Google. domain The target's domain, such as `de` for amazon.de. context `key` and `value` items that the API reads from the `context` list. storage_type The Cloud Storage type that uploads the result, with Push-Pull only. Only `gcs` has a live upload test. storage_url The bucket path that Cloud Storage uploads to. A path that ends in `.{{ extension }}` names each job's object, so it raises without `{{ job_id }}`: jobs that share a name lose their uploads and still bill. `repr`, validation errors and `dry_run` show its credentials as `redacted:redacted`, as the API does. The API returns a free 400 for a raw `/`, `?` or `#` in the secret, and accepts it percent-encoded. A document that is not valid JSON fails before any `Payload` code runs, so only `Payload.model_validate_json` redacts that error. A caller's `TypeAdapter` or model that holds a `Payload` keeps the whole document in `errors()` and `json()`, credentials included. parsing_instructions The instructions of a custom parser, which need `parse`. A wrong `_args` shape, or a regex that Python's `re` cannot compile, raises, because the API bills it with a null field. browser_instructions The browser actions to run on the page, which need `render`. An instruction after `fetch_resource`, or a `filter` that Python's `re` cannot compile, raises, because the API returns 500 for it on every attempt. extra Keys in the API's shape, which oxy merges into the body. Its `context` items follow the typed ones, and a key set both here and as a field raises. It also carries a value that an out-of-date `Literal` rejects. Universal(*, source: Literal['universal'] = 'universal', url: str, render: Optional[Literal['html', 'png', '']] = None, user_agent_type: Optional[Literal['desktop', 'mobile', 'mobile_android', 'mobile_ios', 'tablet', 'tablet_android', 'tablet_ios']] = None, callback_url: str | None = None, parse: bool | None = None, markdown: bool | None = None, xhr: bool | None = None, parser_preset: str | None = None, content_encoding: Optional[Literal['base64', 'utf-8']] = None, client_notes: str | None = None, aggregate_name: str | None = None, geo_location: str | None = None, context: list[oxyscraper._payloads._ContextItem] | None = None, storage_type: Optional[Literal['gcs', 's3', 'tos', 's3_compatible']] = None, storage_url: str | None = None, parsing_instructions: Optional[Annotated[ParsingInstructions, PlainValidator(func=, json_schema_input_type=Any), WithJsonSchema(json_schema={}, mode=None)]] = None, browser_instructions: list[typing.Annotated[oxyscraper._payloads._Click | oxyscraper._payloads._Input | oxyscraper._payloads._Scroll | oxyscraper._payloads._Wait | oxyscraper._payloads._FetchResource, FieldInfo(annotation=NoneType, required=True, discriminator='type')]] | None = None, extra: dict[str, typing.Any] = {}, force_headers: bool | None = None, force_cookies: bool | None = None, successful_status_codes: list[int] | None = None, follow_redirects: bool | None = None, cookies: list[oxyscraper._payloads._Cookie] | None = None, headers: dict[str, str] | None = None, session_id: str | None = None, http_method: Optional[Literal['get', 'post', 'options']] = None, content: str | None = None, store_id: str | None = None) -> None A `universal` job, which scrapes any URL. The model takes no `domain`, `locale`, `start_page`, `pages` or `limit`, because no docs page names them for `universal`. `extra` passes them. [What a live test shows about `universal` and the instruction parameters](https://github.com/ozanozbeker/oxyscraper/blob/main/docs/research/live-universal.md) records the run that checked it. Attributes ---------- url The page to scrape. A page with an empty body faults the job, whatever its status. geo_location A country name or an ISO 3166-1 alpha-2 code, such as `Germany` or `DE`. The API accepts any string, and a value it does not know, such as `de`, has no effect and still bills. user_agent_type The device of the job's user agent. A `desktop_*` value raises, because it draws from the same agents as `desktop`. force_headers Sends `headers` to the site. force_cookies Sends `cookies` to the site. successful_status_codes More status codes that end the job `done`, such as 503. The API rejects a 3xx code for free. follow_redirects `False` faults a job whose page redirects. A chain of more than 10 redirects faults the job either way. cookies The cookies that the site receives. They raise without `force_cookies`, because the site then receives none and the job still bills. headers The headers that the site receives. They raise without `force_headers`, because the site then receives none and the job still bills. A `User-Agent` header never replaces Oxylabs' own. session_id Jobs that share an ID share an exit IP, for 100 jobs or 25 minutes after the first. http_method `post` sends `content` as the request body. content The request body in Base64, which the API rejects for free in any other encoding. store_id A Home Depot store ID. AsyncRun(jobs: 'MemoryObjectReceiveStream[Job]', state: '_RunState') -> 'None' The jobs of one `AsyncSession.stream` call, each as it finishes. AsyncSession(*, username: 'str', password: 'str', transport: 'httpx2.AsyncBaseTransport | None' = None, retry_limit: 'float' = 600.0, pending_limit: 'float | None' = 600.0, progress_interval: 'float | None' = 10.0) -> 'None' Submit payloads and fetch their jobs inside an `async with` block, on asyncio or trio. The block holds the task group that submits and checks while the caller's loop body runs. Leaving it stops every unfinished run. Parameters ---------- username The API user's username. password The API user's password. transport The transport of every request, such as a `FakeOxylabs`. `None` uses the fake of the innermost open `with FakeOxylabs()` block, or else Oxylabs over HTTP/2. retry_limit The seconds after a request's first failure until it stops retrying. pending_limit The seconds after a job's acceptance until oxy stops checking it. `None` waits without limit. progress_interval The seconds between progress lines. `None` turns the lines off. Raises ------ ValueError If `username` or `password` is empty. Run(jobs: 'Iterator[Job]', progress: 'Callable[[], Progress]') -> 'None' The jobs of one run, each as it finishes. Iteration yields each job once, so `all`, `one` and `partitions` return only the jobs it has not yet yielded. After the last job, it raises `IncompleteRunError` if a payload ended with no done or faulted job or an upload failed, and raises it again on each later call. Session(*, username: 'str', password: 'str', transport: 'httpx2.AsyncBaseTransport | None' = None, retry_limit: 'float' = 600.0, pending_limit: 'float | None' = 600.0, progress_interval: 'float | None' = 10.0) -> 'None' Submit payloads and fetch their jobs from sync code, inside a `with` block. It takes the arguments of `AsyncSession`, and runs one through an anyio blocking portal. The portal's thread stays open only inside the block, because an open portal blocks the interpreter's exit. ## Dataclasses Data-holding classes DryRun(*, jobs: 'list[dict[str, Any]]', max_results: 'int') -> None The jobs a run would submit. Attributes ---------- jobs Each job's body, with the credentials in `storage_url` replaced by `redacted:redacted`. max_results The most results the jobs can bill: the sum of their `pages`, read as the API reads them, with 1 for a job without a positive one. Rejected and faulted jobs bill nothing, and a source that ignores `pages` bills 1, so a run can bill less. Job(*, id: 'str', status: '_Status', source: 'str', input: 'str', created_at: 'datetime', finished_at: 'datetime | None', payload: 'Payload | None', results: 'list[Result]', data: 'dict[str, Any]', upload: 'Upload | None') -> None One job and its results. Attributes ---------- id The job's ID. status `pending` until the job finishes. source The job's source. input The value of the job's input key, such as its `url` or `query`. created_at When the API accepted the job, in UTC. finished_at When the job finished, in UTC, or `None` while it is pending. It comes from the results, because a Realtime job object keeps `updated_at` equal to `created_at`. payload The payload that created the job, or `None` for a job from `get`. results One entry per page and output type. A pending job has none, and so has a job from `get` whose results expired. data The job object as the API returned it, so a field that oxy does not type stays readable. upload The Cloud Storage upload, or `None` without `storage_type`. Progress(*, payloads: 'int', unsubmitted: 'int', pending: 'int', done: 'int', faulted: 'int', rejected: 'int', unfetched: 'int', written: 'int', uploaded: 'int', unuploaded: 'int', retries: 'int', elapsed: 'timedelta') -> None The number of a run's payloads in each state, at one moment. The six states from `unsubmitted` to `unfetched` sum to `payloads`. After the run's last job, the snapshot stops changing and is the run's summary. Attributes ---------- payloads The run's payloads. unsubmitted The payloads not yet sent, including those that oxy holds back. pending The payloads whose job is pending. done The payloads whose job is done. faulted The payloads whose job faulted. rejected The payloads that the API rejected, so no job exists. unfetched The payloads whose job oxy stopped checking. written The done jobs that oxy wrote to the destination. uploaded The jobs whose Cloud Storage upload succeeded. unuploaded The jobs whose Cloud Storage upload failed. retries The requests that the run sent again. elapsed The time since the run started, or the run's length once it ends. Rejection(*, payload: 'Payload', status_code: 'int', message: 'str', trace_id: 'str | None') -> None The error that the API returned for a payload instead of a job, so nothing billed. Attributes ---------- payload The rejected payload. status_code The response's status code. message The response's message, as `OxylabsError.message` reads it. trace_id The response's trace ID. Result(*, page: 'int', type: '_OutputType', status_code: 'int', content: '_Content', created_at: 'datetime', updated_at: 'datetime', data: 'dict[str, Any]') -> None One page of a job's results, in one output type. Attributes ---------- page The page number. type The output type. status_code The target's status code. A faulted job's entry holds 613 or 400, and a faulted job can have no entry. content The content, with `png` decoded from Base64 to bytes. created_at When the job started, in UTC. updated_at When the result finished, in UTC. data The result entry as the API returned it. Upload(*, storage_url: 'str', code: 'int | None', message: 'str | None') -> None A job's Cloud Storage upload. Attributes ---------- storage_url The object's path, which the API resolved from the payload's `storage_url`. code The code of the first entry in the job's `statuses`, where 13000 means success. `None` means no entry appeared before the pending limit, or the job object has no `statuses`, which holds for the sources whose job object is the payload alone. oxy cannot check those uploads, so it counts them neither as uploaded nor as unuploaded. message That entry's message. ## Exceptions Exception classes IncompleteRunError(*, rejections: 'list[Rejection]', unsubmitted: 'list[Payload]', unfetched: 'list[Job]', unuploaded: 'list[Job]') -> 'None' The error a run raises after its last job, when a payload ended with no done or faulted job, an upload failed, or a failure stopped the run. Its `__cause__` is the `OxylabsError` or the failed write that stopped the run, if one did. Attributes ---------- rejections One rejection per rejected payload. unsubmitted The payloads that a stop left unsent. unfetched The jobs that oxy stopped checking, each pending and with no results. `get` fetches one later. unuploaded The jobs whose Cloud Storage upload failed. jobs The jobs that `all` collected before it raised, and empty for any other call. OxylabsError(*, status_code: 'int | None', message: 'str', trace_id: 'str | None' = None) -> 'None' One error response from the API, or a network failure that lasted past the retry limit. It chains the httpx2 exception. Attributes ---------- status_code The response's status code, or `None` for a network failure. message The body's `message`, or its `errors` joined when it has none, or else the reason phrase. trace_id The ID that Oxylabs support asks for, or `None` when the response names none. ## Functions Public functions dry_run(payloads: 'Payload | Iterable[Payload]') -> 'DryRun' List the jobs that a run of `payloads` would submit, and the most results they can bill. It sends no request and needs no credentials. oxy never reads the account's remaining results, so a caller who wants a spending limit compares `max_results` with their own. Examples -------- ```python import oxyscraper as oxy report = oxy.dry_run( oxy.Payload(source="universal", url="https://example.com", pages=2) ) print(report.job_count, report.max_results) # 1 2 ``` ## Constants Module-level constants and data SOURCES Built-in immutable sequence. If no argument is given, the constructor returns an empty tuple. If iterable is specified the tuple is initialized from iterable's items. If the argument is a tuple, the return value is the same object. Annotated(*args, **kwargs) Runtime representation of an annotated type. At its core 'Annotated[t, dec1, dec2, ...]' is an alias for the type 't' with extra annotations. The alias behaves like a normal typing alias. Instantiating is the same as instantiating the underlying type; binding it to types is also the same. The metadata itself is stored in a '__metadata__' attribute as a tuple. ParsingFunction Represent a PEP 604 union type E.g. for int | str ParsingInstructions() Create named, parameterized type aliases. This provides a backport of the new `type` statement in Python 3.12: type ListOrSet[T] = list[T] | set[T] is equivalent to: T = TypeVar("T") ListOrSet = TypeAliasType("ListOrSet", list[T] | set[T], type_params=(T,)) The name ListOrSet can then be used as an alias for the type it refers to. The type_params argument should contain all the type parameters used in the value of the type alias. If the alias is not generic, this argument is omitted. Static type checkers should only support type aliases declared using TypeAliasType that follow these rules: - The first argument (the name) must be a string literal. - The TypeAliasType instance must be immediately assigned to a variable of the same name. (For example, 'X = TypeAliasType("Y", int)' is invalid, as is 'X, Y = TypeAliasType("X", int), TypeAliasType("Y", int)'). ---------------------------------------------------------------------- This is the User Guide documentation for the package. ---------------------------------------------------------------------- ## Guide ### Payloads {{< include _includes/setup.qmd >}} A payload holds the parameters of one job. A session sends it as the body of a submission, and the API creates one job from it. ## Install ```sh uv add oxyscraper ``` The library installs no CLI packages. [The command line](cli.qmd) shows how to install the `oxy` command. ## Build a payload Each source that a billed run has checked has its own model, and `oxy.SOURCES` lists them: ```{python} import oxyscraper as oxy [model.__name__ for model in oxy.SOURCES] ``` A model fixes `source` and types each parameter that its source takes. Each allowed value is a `Literal`, so an editor lists the values and a type checker flags a misspelt one. `model_dump()` returns the body that oxy sends: ```{python} payload = oxy.AmazonSearch( query="standing desk", domain="de", sort_by="price_low_to_high", pages=2 ) payload.model_dump() ``` The API reads `sort_by` from the `context` list, so the model moves it there. You pass every parameter as a plain keyword, and never build `context` by hand. An unset field stays out of the body, so the API applies its own default. A job's `data` then records the value the API chose. ## Sources without a model `oxy.Payload` takes any source and any parameter. Each keyword that it does not type goes into the body as it is: ```{python} oxy.Payload(source="walmart_product", product_id="436012154").model_dump() ``` `Payload` types the parameters that keep one name, type and value set on every source that takes them, such as `render`, `parse`, `pages` and `storage_url`. `geo_location`, `locale` and `domain` stay `str`, because their values vary by source. Every payload sets exactly one input key, such as `query`, `url` or `product_id`, to a non-empty string. ## Mistakes that bill A model raises for each mistake that a live run showed billing. A misspelt keyword is the most common, because the API bills a job that ignores it: ```{python} import pydantic try: oxy.AmazonProduct(query="B07FZ8S74R", domian="de") except pydantic.ValidationError as error: print(error) ``` An ASIN longer than 10 characters raises too, because the API bills a 404 page for it. A model leaves each mistake that the API rejects for free to the API. The run reports such a payload as a rejection, which [Failures](failures.qmd) describes. Each model's reference page lists its source's rules and caveats, such as which `geo_location` values may bill on [`AmazonSearch`](../reference/AmazonSearch.qmd). ## Parameters that oxy does not type `extra` passes keys in the API's own shape. The model merges them into the body, and appends their `context` items after the typed ones. `extra` also carries a value that oxy's `Literal` does not list yet, such as a new `sort_by` order: ```{python} oxy.AmazonSearch( query="standing desk", extra={"context": [{"key": "sort_by", "value": "newest_arrivals"}]}, ).model_dump() ``` A key set both as a field and in `extra` raises. ## Instructions `parsing_instructions` and `browser_instructions` are typed too: ```{python} oxy.Universal( url="https://sandbox.oxylabs.io/products", render="html", parse=True, browser_instructions=[ {"type": "scroll_to_bottom"}, {"type": "wait", "wait_time_s": 2}, ], parsing_instructions={ "titles": {"_fns": [{"_fn": "xpath", "_args": ["//h4/text()"]}]}, }, ) ``` A wrong `_args` shape raises, because the API bills the job and returns a null field. An instruction after `fetch_resource` raises, because the API returns 500 for it on every attempt. ## Check a run's size `oxy.dry_run` lists the body of each job and the most results the jobs can bill. It sends no request and needs no credentials: ```{python} report = oxy.dry_run( [oxy.AmazonSearch(query=query, pages=3) for query in ["standing desk", "desk lamp"]] ) print(report.job_count, report.max_results) ``` `max_results` is an upper bound. A rejected or faulted job bills nothing, and a source that ignores `pages` bills one result. The dry run never reads the account's remaining results, so compare `max_results` with your own limit. ### Sessions and runs {{< include _includes/setup.qmd >}} A session submits payloads and checks their jobs in the background, for as long as its `with` block runs. Each `execute` call starts one run. ## Run payloads `Session` takes the username and password of a Web Scraper API (Classic) user as arguments. The library reads no environment variable and no `.env` file, so your code chooses where the credentials come from: ```{python} import os import oxyscraper as oxy USERNAME = os.environ["OXY_WSA_USERNAME"] PASSWORD = os.environ["OXY_WSA_PASSWORD"] payloads = [ oxy.Universal(url=f"https://sandbox.oxylabs.io/products/{number}") for number in range(1, 4) ] with oxy.Session(username=USERNAME, password=PASSWORD) as session: run = session.execute(payloads) for job in run: print(job.id, job.status, job.input) ``` An empty username or password raises `ValueError` before any request. The Oxylabs docs lead with their newer Web API, and a new account may hold only a Web API key. The Classic API returns 401 for that key, and the run stops as [Errors that stop a run](failures.qmd#errors-that-stop-a-run) describes. `execute` starts submitting at once and returns a `Run`. The run yields each job as it finishes, so the loop body handles early jobs while oxy submits and checks later ones. Leaving the `with` block stops every run that has not finished. A run also has three methods, and each returns only the jobs that the run has not yet yielded: - `all()` waits for every job and returns them in a list. - `one()` returns the run's only job, and raises unless the run holds exactly one. - `partitions(size)` yields lists of `size` jobs as they finish, for a bulk write. ## Jobs ```{python} with oxy.Session(username=USERNAME, password=PASSWORD) as session: job = session.execute( oxy.Universal(url="https://sandbox.oxylabs.io/products/1") ).one() job ``` `job.content` returns the content of the job's only result, and raises if the job has more or fewer than one. `job.results` lists one `Result` per page and output type, with `png` content decoded to bytes. `job.data` holds the job object as the API returned it, so a field that oxy does not type stays readable. `output_types` asks for several output types at once, so one fetch returns both raw and parsed content: ```{python} with oxy.Session(username=USERNAME, password=PASSWORD) as session: run = session.execute( oxy.AmazonProduct(query="B07FZ8S74R", parse=True), output_types=["raw", "parsed"], ) job = run.one() [(result.type, type(result.content).__name__) for result in job.results] ``` ## Progress `run.progress` returns a frozen `Progress`, which counts the run's payloads in each state. After the last job it stops changing, so it is also the run's summary: ```{python} print(run.progress) ``` The six states, `unsubmitted`, `pending`, `done`, `faulted`, `rejected` and `unfetched`, sum to `payloads`. `retries` counts every request that oxy sent again, so it rises during an outage. Each field is an `int` or a `timedelta`, so you can record them as asset metadata. ## Realtime A run uses Push-Pull unless you pass `realtime=True`. A session never switches the integration method, and both methods yield the same `Job`: ```{python} with oxy.Session(username=USERNAME, password=PASSWORD) as session: run = session.execute( oxy.Universal(url="https://sandbox.oxylabs.io/products/1"), realtime=True ) print(run.one().status) ``` Realtime returns 408 for a job that runs 150 seconds or longer, and oxy reports that payload as a rejection that points to Push-Pull. ## Fetch a job by ID `get` returns a job as it stands, at once: ```{python} with oxy.Session(username=USERNAME, password=PASSWORD) as session: fetched = session.get(job.id) fetched.status, fetched.payload ``` A pending job comes back with `status="pending"` and no results. So does a finished job whose results expired, but with its final status. `payload` is `None` for a job from `get`. The API returns 404 for the ID of a Realtime job, so `get` fetches Push-Pull jobs only. ## Async code `AsyncSession` takes the same arguments, and runs on asyncio and trio. `stream` returns an `AsyncRun`, which yields each job as it finishes: ```{python} async def scrape() -> list[str]: async with oxy.AsyncSession(username=USERNAME, password=PASSWORD) as session: run = await session.stream(payloads) return [job.id async for job in run] await scrape() ``` A notebook awaits it at the top level, as above, and a script calls `asyncio.run(scrape())`. `AsyncSession.execute` awaits every job and returns a `Run`. ## Scheduling A run needs no setting for its pace: - It sends payloads that share every parameter except the input as one batch, so 8,600 payloads take about 172 submissions. - It paces submissions from the rate-limit headers of each response, so a run stays under the plan's limit. - It checks each job 1 second after the API accepts it, then every second until the job is 10 seconds old, then every 5 seconds. - It stops checking a job that is still pending after `pending_limit` seconds, 600 by default, and reports it as unfetched. - It holds at most 100 finished jobs for your loop, and pauses while they wait, so a slow loop keeps memory bounded. Two runs on one session share its rate-limit budgets and its limit of 100 requests at once. Processes share no limit. Each process slows down as the headers show the account's remaining requests drop, and retries each 429. To run one run at a time across processes, limit concurrency outside oxy, such as with a Dagster pool. Oxylabs has not said whether API users under one account share one limit, so a second API user may add no throughput. ## Logging The library writes to the `oxyscraper` logger and attaches no handler, so your logging config sets where its lines go. WARNING covers each event that loses a result or may cost money. INFO covers the start of a run, a progress line every `progress_interval` seconds, the end line and the run log's path. The library writes no ERROR line, because it raises instead. httpx2 logs each request at INFO, and a run that checks 8,600 jobs writes thousands of those lines. Set the `httpx2` logger to WARNING: ```{python} import logging logging.basicConfig(level=logging.INFO, format="%(levelname)s %(name)s: %(message)s") logging.getLogger("httpx2").setLevel(logging.WARNING) with oxy.Session(username=USERNAME, password=PASSWORD) as session: session.execute(payloads).all() ``` ### Failures {{< include _includes/setup.qmd >}} ```{python} #| include: false import math from oxyscraper.testing import Outcome faulted_once: set[str] = set() def outcome(payload: dict[str, str]) -> Outcome: url = payload["url"] if url.endswith("/2") and url not in faulted_once: faulted_once.add(url) return Outcome(status="faulted") if url.endswith("/3"): return Outcome(after=math.inf) if url.endswith("/missing"): return Outcome(status_code=404) return Outcome() fake = ExitStack().enter_context(FakeOxylabs(outcome)) ``` A run reports what happened to every payload. It yields each done or faulted job, and raises one `IncompleteRunError` after its last job if any payload ended without one. ## A run with failures Here Oxylabs faults the second job, the third job never finishes, and the API rejects the fourth URL for free. A short `pending_limit` makes oxy stop checking the third job after 5 seconds instead of 10 minutes: ```{python} import os import oxyscraper as oxy USERNAME = os.environ["OXY_WSA_USERNAME"] PASSWORD = os.environ["OXY_WSA_PASSWORD"] payloads = [ oxy.Universal(url=f"https://sandbox.oxylabs.io/products/{number}") for number in range(1, 4) ] payloads.append(oxy.Universal(url="https://10.0.0.1/")) faulted = [] with oxy.Session(username=USERNAME, password=PASSWORD, pending_limit=5) as session: run = session.execute(payloads) try: for job in run: print(job.status, job.input) if job.status == "faulted": faulted.append(job.payload) except oxy.IncompleteRunError as error: print(error) for rejection in error.rejections: print(rejection.status_code, rejection.message, rejection.payload.url) for job in error.unfetched: print(job.id, job.status, job.input) ``` ## Faulted jobs A faulted job is one that Oxylabs could not complete, even after retrying it, and it bills nothing. The run yields it with `status="faulted"`, and a run whose only problem is faulted jobs ends without raising. The run never resubmits a faulted job, so you choose whether a resubmission is worth the wait: ```{python} with oxy.Session(username=USERNAME, password=PASSWORD) as session: for job in session.execute(faulted): print(job.status, job.input) ``` ## The lists of `IncompleteRunError` The error carries one list for each way a payload can end without a done or faulted job: - `rejections` holds one `Rejection` per payload that the API rejected, with its `status_code`, `message` and `trace_id`. A 400, a 422, an entry in a batch's `errors` and a Realtime 408 each reject their payload, and the run carries on. A payload rejected inside a batch carries the batch's 202, as the fourth URL above does. - `unsubmitted` holds the payloads that a stop left unsent. - `unfetched` holds the jobs that oxy stopped checking, each pending and with no results. `session.get` fetches one later. - `unuploaded` holds the jobs whose Cloud Storage upload failed. - `jobs` holds the jobs that `all()` collected before it raised. At the end, `run.progress` counts the same payloads: ```{python} print(run.progress) ``` ## Errors that stop a run A 401, a 403, the domain throttle, a submission out of retries and a failed write to the destination each stop submission. The run still checks the jobs that the API accepted and yields them, then raises `IncompleteRunError`. Its `__cause__` is the `OxylabsError` that stopped the run, and the payloads it never sent are in `unsubmitted`. `OxylabsError` carries the `status_code`, `message` and `trace_id` of one error response, which Oxylabs support asks for. `get` raises it directly: ```{python} try: with oxy.Session(username=USERNAME, password=PASSWORD) as session: session.get("7500000000000000999") except oxy.OxylabsError as error: print(error.status_code, error.message) ``` ## Retries Each request retries after a 429, any 5xx or a network error. Each wait is random, between 0 and a ceiling that starts at 1 second and doubles up to 30 seconds. A request stops retrying `retry_limit` seconds after its first failure, 600 by default. A check that runs out of retries moves only its own job into `unfetched`. The API has no idempotency key, so a submission retried after a 5xx or a read error may create a second job that bills. Each such retry writes a WARNING line. ## Results that are not failures A done job whose page returned a 4xx stays done and bills. Each `Result` holds the target's `status_code`, so check it when a 404 matters: ```{python} with oxy.Session(username=USERNAME, password=PASSWORD) as session: job = session.execute( oxy.Universal(url="https://sandbox.oxylabs.io/products/missing") ).one() job.status, job.results[0].status_code ``` ## Stopping a run Ctrl+C, a `break` out of the loop or leaving the `with` block stops the run. Each submission that oxy already sent finishes, so the run log records the ID of every job that may bill. A second Ctrl+C stops at once, without the run log. ### Destinations and run logs {{< include _includes/setup.qmd >}} A destination receives each done job as a file, and a run log records what happened to each payload. Together they let a large run write its results to disk or a bucket, and leave a record to recover from. ## Write each job to a destination `destination` names a folder. The run writes each done job there as `.json` before it yields the job, so a yielded job's file always exists: ```{python} import os import oxyscraper as oxy USERNAME = os.environ["OXY_WSA_USERNAME"] PASSWORD = os.environ["OXY_WSA_PASSWORD"] payloads = [ oxy.Universal(url=f"https://sandbox.oxylabs.io/products/{number}") for number in range(1, 4) ] with oxy.Session(username=USERNAME, password=PASSWORD) as session: run = session.execute(payloads, destination="results") for _ in run: pass print(run.progress) sorted(os.listdir("results")) ``` Drain such a run with a bare loop, as above. `all()` returns every job with its content, so it holds the whole run in memory. Each file holds the API's body unchanged, as one line of JSON. Only done jobs get a file, so every file in a destination is a billed result. A run never deletes a file, and replaces a file of the same name, so you choose what to clear between runs. `destination` takes three kinds of value: - A local path, which oxy creates. - A URL such as `s3://bucket/path`, `gs://bucket/path` or `az://container/path`, which reads the same `AWS_*`, `GOOGLE_*` and `AZURE_*` variables as Polars. - An obstore store, for explicit credentials, a credential provider, an S3 region or a retry config. The run lists the destination before its first submission, so a missing bucket or a wrong credential raises before anything bills. A failed write stops submission, and the run yields the remaining jobs unwritten before it raises. ## Read a destination with Polars `pl.scan_ndjson` reads the folder. Pass a schema, because the values in `job.context` mix lists and scalars, and Polars cannot infer one type for them: ```{python} import polars as pl schema = { "results": pl.List( pl.Struct({"url": pl.String, "status_code": pl.Int64, "content": pl.String}) ), "job": pl.Struct({"id": pl.String, "source": pl.String, "status": pl.String}), } ( pl.scan_ndjson("results/*.json", schema=schema) .explode("results", empty_as_null=False) .select( pl.col("job").struct.field("id"), pl.col("results").struct.field("url", "status_code"), ) .collect() ) ``` A schema also limits the scan to the fields it names. `content` is a string for `raw` and `markdown` results, and an object for `parsed` ones, so give it the type of the output type you scan. ## Write a run log ```{python} #| include: false from oxyscraper.testing import Outcome faulted_once: set[str] = set() def outcome(payload: dict[str, str]) -> Outcome: if payload["url"].endswith("/2") and payload["url"] not in faulted_once: faulted_once.add(payload["url"]) return Outcome(status="faulted") return Outcome() fake = ExitStack().enter_context(FakeOxylabs(outcome)) # The first check fails, so its job ends unfetched. fake.fail(404, on="results") ``` `run_log` names a folder, and takes the same values as `destination`. Each run writes one file there, named by the run's start in UTC, such as `20260929T170412.345Z.jsonl`: ```{python} with oxy.Session(username=USERNAME, password=PASSWORD) as session: run = session.execute(payloads, destination="results", run_log="logs") try: for _ in run: pass except oxy.IncompleteRunError as error: print(error) ``` The file holds one line per payload, in input order, and each line has the same five keys: ```{python} import json from pathlib import Path log = max(Path("logs").glob("*.jsonl")) lines = [json.loads(line) for line in log.read_text().splitlines()] lines ``` ```{python} #| include: false assert [line["state"] for line in lines] == ["unfetched", "faulted", "done"], lines ``` - `state` is `done`, `faulted`, `rejected`, `unsubmitted` or `unfetched`. - `id` is the job's ID, or `null` when no job exists. - `payload` is the body that oxy sent. - `error` holds the error behind a rejected, unsubmitted or unfetched payload. - `upload` holds the Cloud Storage entry's `code` and `message`. The run writes an empty file before its first submission, so a folder it cannot write to raises before anything bills. It replaces the file when the run ends, raises or stops, even after Ctrl+C. An empty run log therefore means a process killed outright or a second Ctrl+C, and jobs may have billed without a record. ## Recover a run Nothing resumes a run, and a second call with the same payloads submits each one again. The run log lists what to redo instead. Resubmit the faulted payloads, and fetch the unfetched jobs by ID, so you pay for neither twice: ```{python} faulted = [ oxy.Payload.model_validate(line["payload"]) for line in lines if line["state"] == "faulted" ] unfetched = [line["id"] for line in lines if line["state"] == "unfetched"] with oxy.Session(username=USERNAME, password=PASSWORD) as session: for job in session.execute(faulted, destination="results"): print(job.status, job.input) for job_id in unfetched: job = session.get(job_id) print(job.status, job.input) ``` `get` writes nothing to a destination, and [`oxy get -d`](cli.qmd) does. The run log replaces the credentials in a `storage_url` with `redacted:redacted`. So a `tos` or `s3_compatible` payload from a run log needs its secret again before a resubmission. ### Cloud Storage {{< include _includes/setup.qmd >}} ```{python} #| include: false from oxyscraper.testing import Outcome def outcome(payload: dict[str, str]) -> Outcome: if payload.get("storage_url", "").startswith("missing-bucket/"): return Outcome(upload=13102) return Outcome() fake = ExitStack().enter_context(FakeOxylabs(outcome)) ``` Cloud Storage makes Oxylabs upload each job's result to your bucket, so the results never pass through your machine. Grant Oxylabs write access to the bucket first, as [Cloud Storage](https://developers.oxylabs.io/products/web-scraper-api/features/result-processing-and-storage/cloud-storage) describes. ## Upload results Set `storage_type` and `storage_url` on each payload. A run then waits for each job's upload, and logs which payload it checks first: ```{python} import logging import os import oxyscraper as oxy logging.basicConfig(level=logging.INFO, format="%(levelname)s %(message)s") logging.getLogger("httpx2").setLevel(logging.WARNING) USERNAME = os.environ["OXY_WSA_USERNAME"] PASSWORD = os.environ["OXY_WSA_PASSWORD"] payloads = [ oxy.Universal( url=f"https://sandbox.oxylabs.io/products/{number}", storage_type="gcs", storage_url="my-bucket/products", ) for number in range(1, 4) ] with oxy.Session(username=USERNAME, password=PASSWORD) as session: for job in session.execute(payloads): print(job.upload) ``` A `storage_url` that names a folder gets one object per job, named by the job's ID. One that ends in `.{{ extension }}` names each object, so it raises without `{{ job_id }}`, because jobs that share a name lose their uploads and still bill. `job.upload` holds the object's path, as the API resolved it, and the code of the upload's entry in the job's `statuses`. A run downloads nothing for such a job, so its `results` stay empty. So `destination` and `realtime=True` raise `ValueError` with a payload that sets `storage_type`, before any request. ## Check the bucket first By default, oxy submits the first payload of each `storage_url` alone, and holds the rest until its upload succeeds. So a wrong bucket bills one job instead of the whole run: ```{python} payloads = [ oxy.Universal( url=f"https://sandbox.oxylabs.io/products/{number}", storage_type="gcs", storage_url="missing-bucket/products", ) for number in range(1, 4) ] with oxy.Session(username=USERNAME, password=PASSWORD) as session: run = session.execute(payloads) try: run.all() except oxy.IncompleteRunError as error: print(error) print(error.unuploaded[0].upload) ``` `check_storage=False` submits every payload at once, and oxy still checks each upload. A failed upload leaves the job `done`, so oxy reads each job's upload code, and counts any code but 13000 as a failed upload. It reads the code, not the message, because the API's messages differ from its docs. A job with no entry by the pending limit counts as a failed upload too. A run never retries an upload, and lists each failed one in `IncompleteRunError.unuploaded`. Some sources return a job object without `statuses`, so oxy cannot see their uploads. Their `upload.code` is `None`, and oxy counts them neither as uploaded nor as unuploaded. ## Storage types `gcs` is the only storage type with a live upload test. No test has uploaded to `s3`, `s3_compatible` or `tos`, so keep `check_storage` on for them. A `tos` or `s3_compatible` URL carries an access key and a secret. In `repr`, validation errors, log lines, the dry run and the run log, oxy replaces them with `redacted:redacted`, as the API does: ```{python} oxy.Universal( url="https://sandbox.oxylabs.io/products/1", storage_type="s3_compatible", storage_url="https://KEY_ID:SECRET@s3.example.com/my-bucket/products", ) ``` The body that oxy sends keeps the secret. A payload that you rebuild from a run log does not, so set the secret again before you resubmit it. ### Testing with the fake {{< include _includes/setup.qmd >}} `oxyscraper.testing.FakeOxylabs` is an httpx2 transport that returns the Web Scraper API's responses, so a test needs no credentials and bills nothing. It applies the API's free checks and rate limits, and keeps its jobs between sessions. Every example on this site runs against it when the site builds. The fake's behaviour is public API under SemVer, so an oxy upgrade never changes your tests' results without a changelog entry. ## Switch it on for a suite `with FakeOxylabs():` points every session built inside the block at the fake, even one that another thread builds. So code that builds its own `Session`, such as a Dagster resource, runs against it, and an autouse fixture guards a whole suite: ```{python} from collections.abc import Iterator import pytest from oxyscraper.testing import FakeOxylabs @pytest.fixture(autouse=True) def fake() -> Iterator[FakeOxylabs]: with FakeOxylabs() as fake: yield fake ``` Here is the code under test, which builds its own session: ```{python} import oxyscraper as oxy def scrape(urls: list[str]) -> list[oxy.Job]: with oxy.Session(username="USERNAME", password="PASSWORD") as session: return session.execute([oxy.Universal(url=url) for url in urls]).all() ``` By default, every job finishes at once, with content that the fake writes from the job's input. `fake.jobs` holds each job object that the fake created, and `fake.requests` each request it received: ```{python} def test_scrape_returns_each_job(fake: FakeOxylabs) -> None: jobs = scrape(["https://sandbox.oxylabs.io/products/1"]) assert [job.status for job in jobs] == ["done"] assert fake.jobs[0]["url"] == "https://sandbox.oxylabs.io/products/1" ``` ```{python} #| include: false with FakeOxylabs() as fake: test_scrape_returns_each_job(fake) ``` An explicit `transport=` on a session wins over the block, and the innermost open block wins over the outer ones. ## Set what each job does An `Outcome` sets what Oxylabs does with a job: its final status, how many seconds it takes, its content, the target's status code and its upload. A function of the payload sets one outcome per job, and its `Rejected` rejects a payload instead. The function receives each payload as the API received it, once per batch value, because a test cannot predict how oxy groups payloads: ```{python} from typing import Any from oxyscraper.testing import Outcome, Rejected def outcome(payload: dict[str, Any]) -> Outcome | Rejected: if payload["url"].endswith("/2"): return Outcome(status="faulted") if payload["url"].endswith("/3"): return Rejected("The hostname cannot be an ip address.") return Outcome(content="Product 1") def test_scrape_keeps_faulted_jobs() -> None: with FakeOxylabs(outcome), pytest.raises(oxy.IncompleteRunError) as caught: scrape( [f"https://sandbox.oxylabs.io/products/{number}" for number in (1, 2, 3)] ) assert sorted(job.status for job in caught.value.jobs) == ["done", "faulted"] assert len(caught.value.rejections) == 1 ``` ```{python} #| include: false test_scrape_keeps_faulted_jobs() ``` `content` takes a string, a parsed object, `bytes`, or a function of the page and the output type. The fake sends `bytes` as Base64, as the API sends `png` content, so a fixture comes back as the same bytes. ## Fail requests `fake.fail` makes the next matching requests return a status, or raise an exception such as `httpx2.ReadTimeout("no answer")`. A failed request creates no job and counts against no limit: ```{python} def test_scrape_retries_an_outage(fake: FakeOxylabs) -> None: fake.fail(503, on="submit") jobs = scrape(["https://sandbox.oxylabs.io/products/1"]) assert [job.status for job in jobs] == ["done"] ``` ```{python} #| include: false with FakeOxylabs() as fake: test_scrape_retries_an_outage(fake) ``` `on` limits the failure to submissions, status checks or results downloads, and `times=None` fails every request from then on. `message` sets the body's message, such as the domain throttle's `Access to example.com has been limited to 1 req/s`. `limit` and `render_limit` set the fake's rate limits, 50 and 13 by default. A submission that does not fit returns 429, as the API does. ## Run slow jobs in no time The fake reads anyio's clock, so trio's `MockClock` runs a 10-minute pending limit in a fraction of a second. With the anyio pytest plugin, a fixture picks the clock: ```{python} from trio.testing import MockClock @pytest.fixture def anyio_backend() -> object: return "trio", {"clock": MockClock(autojump_threshold=0)} ``` ```{python} import math @pytest.mark.anyio async def test_a_stuck_job_ends_unfetched() -> None: payload = oxy.Universal(url="https://sandbox.oxylabs.io/products/1") with FakeOxylabs(Outcome(after=math.inf)): async with oxy.AsyncSession( username="USERNAME", password="PASSWORD" ) as session: run = await session.stream(payload) with pytest.raises(oxy.IncompleteRunError) as caught: await run.all() assert len(caught.value.unfetched) == 1 ``` ```{python} #| include: false import trio trio.run(test_a_stuck_job_ends_unfetched, clock=MockClock(autojump_threshold=0)) ``` ## Test `get` The fake keeps its jobs after a session closes, so a second session fetches what a first one submitted: ```{python} def test_get_fetches_an_earlier_job(fake: FakeOxylabs) -> None: [job] = scrape(["https://sandbox.oxylabs.io/products/1"]) with oxy.Session(username="USERNAME", password="PASSWORD") as session: assert session.get(job.id).status == "done" ``` ```{python} #| include: false with FakeOxylabs() as fake: test_get_fetches_an_earlier_job(fake) ``` ### The command line {{< include _includes/setup.qmd >}} ```{python} #| include: false import json import re import shlex import subprocess from pathlib import Path from rich.console import Console from typer.testing import CliRunner from oxyscraper import _cli from oxyscraper.testing import Outcome # rich renders to the notebook under Jupyter, so the CLI prints to a plain stderr console instead. _cli._console = Console(stderr=True, force_jupyter=False, force_terminal=False) ANSI = re.compile(r"\x1b\[[0-9;]*m") def shell(command: str, code: int = 0) -> None: """Show a shell line and its output, running its last command, `oxy`, in this process, where the fake is.""" before, pipe, args = command.rpartition("| oxy ") if not pipe: args = command.removeprefix("oxy ") args, _, output = args.partition(" > ") args, _, source = args.partition(" < ") stdin = None if source: stdin = Path(source).read_text() elif before: stdin = subprocess.run(before, shell=True, capture_output=True, text=True, check=True).stdout result = CliRunner().invoke(_cli.app, shlex.split(args), input=stdin, prog_name="oxy") assert result.exit_code == code, (result.exit_code, result.output) if output: Path(output).write_text(result.stdout) shown = ANSI.sub("", result.stderr + ("" if output else result.stdout)).rstrip() print(f"```sh\n{command}\n```\n") if shown: print(f"```text\n{shown}\n```\n") # The URLs whose next job faults. faulting: set[str] = set() def outcome(payload: dict[str, str]) -> Outcome: if payload.get("url") in faulting: faulting.remove(payload["url"]) return Outcome(status="faulted") return Outcome() fake = ExitStack().enter_context(FakeOxylabs(outcome)) Path("urls.txt").write_text( "".join(f"https://sandbox.oxylabs.io/products/{number}\n" for number in range(1, 5)) ) ``` The `oxy` command runs jobs from a shell, through the same `Session` that the library uses. ## Install ```sh uv tool install 'oxyscraper[cli]' ``` This installs two commands, `oxy` and `oxyscraper`, which do the same thing. `uv add oxyscraper` installs the library alone, and its commands then print the install command above. `oxy` reads the credentials of a Web Scraper API (Classic) user from `OXY_WSA_USERNAME` and `OXY_WSA_PASSWORD`. No option takes the password, so it never lands in your shell history, and `oxy` loads no `.env` file. ## Run jobs `oxy run SOURCE INPUT...` runs one job per input, and prints each done job to stdout as one line of JSON, the body that the API returned. So a redirect writes a file that `pl.read_ndjson` reads: ```{python} #| echo: false #| output: asis shell("oxy run universal https://sandbox.oxylabs.io/products/1 https://sandbox.oxylabs.io/products/2 > results.ndjson") ``` An input with `://` goes in `url`, and any other input in `query`, so most sources need no flag. `-k` names another input key, such as `-k product_id`. Each parameter that `oxy.Payload` types has its own option, which `oxy run --help` lists with its allowed values. `-p KEY=VALUE` sets any other parameter as a string, and `-p KEY:=JSON` sets a JSON value. `--dry-run` prints each payload and the most results they can bill, and sends nothing: ```{python} #| echo: false #| output: asis shell("oxy run amazon_search 'standing desk' 'desk lamp' --domain de --pages 2 -p sort_by=price_low_to_high --dry-run") ``` `oxy` builds each payload with its source's model when one exists, so a misspelt key or value exits with code 2 before any request: ```{python} #| echo: false #| output: asis shell("oxy run amazon_product B07FZ8S74R -p domian=de", code=2) ``` ## Read stdin Without inputs, `oxy run SOURCE` reads one input per line of stdin. Without a source, it reads one payload per line, as JSON in the API's shape, so a dry run's output feeds a later run: ```{python} #| echo: false #| output: asis shell("oxy run universal --dry-run < urls.txt > payloads.ndjson") shell("oxy run -d results < payloads.ndjson") ``` `-d` writes each done job to `.json` in a folder or a bucket instead of stdout, as `destination` does in [Destinations and run logs](destinations.qmd). A URL such as `gs://bucket/path` names a bucket. ## Recover a run `--run-log DIR` writes a run log, one line per payload with its state, job ID and payload: ```{python} #| echo: false #| output: asis faulting.add("https://sandbox.oxylabs.io/products/3") # The first check fails, so its job ends unfetched. fake.fail(404, on="results") shell("oxy run universal -d results --run-log logs < urls.txt", code=1) log = next(Path("logs").glob("*.jsonl")) states = sorted(json.loads(line)["state"] for line in log.read_text().splitlines()) assert states == ["done", "done", "faulted", "unfetched"], states ``` The run exits with code 1, because a job faulted and oxy stopped checking another. `jq` reads the run log, so two commands recover the run without paying twice. The first resubmits the faulted payloads: ```{python} #| echo: false #| output: asis shell(f"jq -c 'select(.state == \"faulted\") | .payload' {log} | oxy run -d results") ``` The second fetches the unfetched jobs, which may have finished since: ```{python} #| echo: false #| output: asis shell(f"jq -r 'select(.state == \"unfetched\") | .id' {log} | oxy get -d results") ``` `oxy get JOB_ID...` takes IDs as arguments or one per line of stdin. For a job that is still pending, faulted or whose results expired, it writes a warning and exits with code 1. The run log redacts the credentials in a `storage_url`. So a `tos` or `s3_compatible` payload needs its secret again before a resubmission. ## Output stdout holds JSON only, and stderr holds every message, in uv's style. A warning starts with `warning:`, a hint with `hint:` and an error with `error:`. On a terminal, a live line shows the run's progress, and elsewhere a plain progress line prints every 10 seconds. `-v` adds a debug line for each change of state and each retry. The exit code tells the outcome: | Code | Meaning | | --- | --- | | 0 | Every job is done. | | 1 | A job faulted, or the run was incomplete. | | 2 | An option, an input or a payload was invalid, or a credential is missing. | | 130 | Ctrl+C stopped the run. | | 143 | SIGTERM stopped the run, which `oxy` handles as Ctrl+C, so the run log is still written. | ### Upgrading Each breaking release has a section here, written in the pull request that made the break. A new minor version always means a breaking change, and every other release is a patch.