Cloud Storage makes Oxylabs upload each job’s result to your bucket, so the results never pass through your machine. Grant Oxylabs write access to the bucket first, as Cloud Storage describes.
Upload results
Set storage_type and storage_url on each payload. A run then waits for each job’s upload, and logs which payload it checks first:
import logging
import os
import oxyscraper as oxy
logging.basicConfig(level=logging.INFO, format="%(levelname)s %(message)s")
logging.getLogger("httpx2").setLevel(logging.WARNING)
USERNAME = os.environ["OXY_WSA_USERNAME"]
PASSWORD = os.environ["OXY_WSA_PASSWORD"]
payloads = [
oxy.Universal(
url=f"https://sandbox.oxylabs.io/products/{number}",
storage_type="gcs",
storage_url="my-bucket/products",
)
for number in range(1, 4)
]
with oxy.Session(username=USERNAME, password=PASSWORD) as session:
for job in session.execute(payloads):
print(job.upload)
INFO Running 3 payloads with Push-Pull
INFO Checking the upload to my-bucket/products with one job before submitting 2 more payloads
INFO Uploaded the first job to my-bucket/products in 1.00s
Upload(storage_url='my-bucket/products/7500000000000000001.json', code=13000, message='Upload Successful')
INFO Finished 3 payloads in 2.00s: 3 done, 3 uploaded
Upload(storage_url='my-bucket/products/7500000000000000002.json', code=13000, message='Upload Successful')
Upload(storage_url='my-bucket/products/7500000000000000003.json', code=13000, message='Upload Successful')
A storage_url that names a folder gets one object per job, named by the job’s ID. One that ends in .{{ extension }} names each object, so it raises without { job_id }, because jobs that share a name lose their uploads and still bill.
job.upload holds the object’s path, as the API resolved it, and the code of the upload’s entry in the job’s statuses. A run downloads nothing for such a job, so its results stay empty. So destination and realtime=True raise ValueError with a payload that sets storage_type, before any request.
Check the bucket first
By default, oxy submits the first payload of each storage_url alone, and holds the rest until its upload succeeds. So a wrong bucket bills one job instead of the whole run:
payloads = [
oxy.Universal(
url=f"https://sandbox.oxylabs.io/products/{number}",
storage_type="gcs",
storage_url="missing-bucket/products",
)
for number in range(1, 4)
]
with oxy.Session(username=USERNAME, password=PASSWORD) as session:
run = session.execute(payloads)
try:
run.all()
except oxy.IncompleteRunError as error:
print(error)
print(error.unuploaded[0].upload)
INFO Running 3 payloads with Push-Pull
INFO Checking the upload to missing-bucket/products with one job before submitting 2 more payloads
WARNING The upload of job 7500000000000000004 failed with 13102 No such path: universal https://sandbox.oxylabs.io/products/1
WARNING Held back 2 payloads for missing-bucket/products, because the first upload failed
INFO Finished 3 payloads in 1.00s: 1 done, 2 unsubmitted, 1 unuploaded
2 unsubmitted, 1 unuploaded
Upload(storage_url='missing-bucket/products/7500000000000000004.json', code=13102, message='No such path')
check_storage=False submits every payload at once, and oxy still checks each upload.
A failed upload leaves the job done, so oxy reads each job’s upload code, and counts any code but 13000 as a failed upload. It reads the code, not the message, because the API’s messages differ from its docs. A job with no entry by the pending limit counts as a failed upload too. A run never retries an upload, and lists each failed one in IncompleteRunError.unuploaded.
Some sources return a job object without statuses, so oxy cannot see their uploads. Their upload.code is None, and oxy counts them neither as uploaded nor as unuploaded.
Storage types
gcs is the only storage type with a live upload test. No test has uploaded to s3, s3_compatible or tos, so keep check_storage on for them.
A tos or s3_compatible URL carries an access key and a secret. In repr, validation errors, log lines, the dry run and the run log, oxy replaces them with redacted:redacted, as the API does:
oxy.Universal(
url="https://sandbox.oxylabs.io/products/1",
storage_type="s3_compatible",
storage_url="https://KEY_ID:SECRET@s3.example.com/my-bucket/products",
)
Universal(source='universal', url='https://sandbox.oxylabs.io/products/1', storage_type='s3_compatible', storage_url='https://redacted:redacted@s3.example.com/my-bucket/products')
The body that oxy sends keeps the secret. A payload that you rebuild from a run log does not, so set the secret again before you resubmit it.
oxyscraper is not affiliated with or endorsed by Oxylabs. Oxylabs and Oxy are trademarks of Oxylabs.