Digercules for Organizations

Most tools for working with corporate data need you to already know what kind of archive you have. Legal software expects case files. A media catalog expects footage. Corporate data audit tools (usually called data governance — services like BigID or IBM Guardium) can find personal data in files and assess leak risk, but can't tell you what's actually inside a file, in terms of meaning.

Digercules starts from a different question: what's actually in here? Bring an unsorted mass of data — a departed employee's computer, an old case archive, decades of raw footage, a music or literary club's recordings of past evenings, a history of client correspondence — and get, not a file list, but a structured library. For every single file — a topic, a language, a document type, key entities, a short summary. This is done ahead of time, without rush or deadline, before anyone urgently needs it — not at the last minute under pressure for a specific task. The finished library can be fed to any neural network, database, or search system — they take it from there, finding what's needed and building the interface; Digercules doesn't replace that next step, it makes it possible in the first place, by taking on the labor-intensive preparation nobody ever has the resources for.

The technology is already working — real numbers from processed projects:

800,000+
files processed
~85 TB
source volume
66,000+ hrs
of audio (~7.5 years)
1.8 TB
reclaimed via deduplication
3,507 hrs
spent processing
~430
performers identified on video

Over the available connection (~30 Mbit/s, measured) uploading this much data to the cloud would take about 262 days — versus the 96 it actually took to process on-site. Storage would run $84/month, and processing on a comparable GPU would run $150–190/month. And that's without the next step: the data still has to move from storage onto the GPU machine itself for processing — even over a connection three times faster (100 Mbit/s), a batch of a few terabytes takes about 3 days.

How it works

Processing runs either entirely on the client's own hardware, or — if that hardware isn't powerful enough — is delegated to a specific external machine over a closed peer-to-peer channel (not a public cloud): the compute is physically located in the Caucasus and Eastern Europe, the client is always told explicitly which machine and which jurisdiction is doing the processing, and only text/structured results come back — never the source files.

This isn't "send data to the cloud." It's a direct transfer to a specific, named device, and only text and structured results come back — not the source files, which of course aren't lost, deleted, or altered in any way.

Why it matters to you

The savings happen in two places at once: your own time (no manual sorting), and the compute budget of whichever neural network works with this next — it gets an already-processed, compact context instead of a raw unsorted mass.

What it costs and how the work proceeds

Free — but only the preliminary estimate. From a description of your archive or a screenshot of the folder, we'll tell you whether it's realistic to sort and roughly how long it would take. Installing the tool and running it on your archive is the paid step — a one-time license for local processing, with no cap on volume or on time. The exact figure for organizations is discussed individually.

Beyond that, if your own hardware isn't enough. The archive passport from the previous step is the basis for pricing the heavier coordination (the part a CPU-only machine can't handle). The price is set by the actual volume and kind of data found — there's no fixed number or ceiling today; the mechanism for calculating it hasn't been settled yet. For a sense of scale, see the figures earlier on this page: what a comparable job would cost on the open cloud market.

What we need from you. Access to the archive (a drive, a folder, an export) and an answer to what you want done with it.

How the transfer actually happens. For organizations, processing by default never leaves your local network. If your own hardware isn't enough, what goes out isn't the archive — it's a compact working export of it (example: from 2 TB of video recordings, about 2.7 GB), sent directly to one agreed-upon machine, never through a public cloud.

If something goes wrong. The support window for retrying a failed run is 48–72 hours from when you report it; a re-run or processing a second drive is a separate arrangement.

Not yet spelled out on this site (worked out case by case): the intake form, a sample contract, exactly where data gets uploaded, the processing schedule and payment schedule, and the acceptance criteria for the result.

Scenarios by industry and task

Customer data and interaction history

Media, production, sound

Professional document archives

Institutions and media infrastructure

Education and practice

How it works under the hood

Wherever people show up on the record, a separate question gets answered: who's actually in frame. How face identification works — the module's principles and its honest limits.

General product description and company structure — on digercules.com. For partnership inquiries — ryazansky@gmail.com.