Repositories and branches
Track public project identity and default branches.
Synchronize public repositories, archives, history, and large objects with unmetered traffic, 10Gbps+ per-project capacity, and source-specific pools.
Discovery, archives, history, and large objects have different transfer patterns and should use separate queues.
Track public project identity and default branches.
Choose current source or full object history.
Transfer only new or changed objects.
Preserve provenance and public access context.
Separate network transfer from parsing, license processing, and deduplication.
Synchronize repositories and source archives.
Preserve commits, trees, and relationships.
Schedule only new or changed objects.
Build a manifest first, then run recoverable repository or object tasks through the proxy layer.
Record source, repository, branch, time, and license.
Match the protocol to history requirements.
Retry at repository or object granularity.
Handle code cleaning and licensing downstream.
Large AI datasets are constrained first by aggregate throughput and transfer cost. Residential proxies are added only when a workload needs explicit geographic identity.
A successful clone alone does not prove references, objects, archives, or large files are complete.
Validate all capacity assumptions with representative public targets, real response sizes, retry behavior, and the required completion window.
Proxy routing does not replace repository discovery, license analysis, code parsing, or deduplication.
git \
-c http.proxy="http://user:pass@proxy.123proxy.cn:9000" \
clone \
--filter=blob:none \
"https://target.example/public/repository.git"
Validate workload boundaries, useful throughput, retry behavior, and dataset completeness with representative public targets.
No. It applies to lawfully accessible public code hosts, archives, and HTTP objects.
Only with valid customer authorization and credentials; the service does not bypass access controls.
Repository history, archives, large files, and retries create sustained transfer.
Use archives for current source and Git when history or object relationships are required.
No. License analysis, code cleaning, and deduplication remain customer-owned.
No, it is aggregate capacity for one project.
Persist update times, references, or object IDs and schedule only changes.
Store source URLs, timestamps, public state, licenses, and processing rules.
Share code sources, repository count, history depth, tools, and completion requirements.