Original Post
Right, figured I'd open this up to some general brainstorming since I've kind of hit a wall with coming up with my own good ideas. (Hey, it's Friday.)
In a nutshell, I've got over 100,000 individual files that need to be tracked and synchronized across several hundred different client locations, and potentially many more clients than that in the future. Old versions of the files need to be available so that outdated clients can patch their way up to current; this is already supported via a delta system but needs to be integrated into a larger architecture. Essentially, I need to generate a manifest of all the files, update this manifest live when files change on an authoritative server, and distribute the manifest (and accompanying files) on demand to clients as they patch up to current.
My overall plan involves something like this:
- Generate a CRC32 of each filename and use this as a hash index for fast lookups of files
- Fall back to canonical filename if the CRC32 hits a collision
- Serve files and the manifest using a custom HTTP server with a virtual file mapping from URIs to the file system/manifest metadata
- Allow serving patches of the manifest given a manifest version ID and a target version ID
This should permit the manifest to grow very large (i.e. hold metadata for all of our files) without requiring that it be distributed in its entirety every time a client needs to patch. A client simply first patches its manifest, then uses the delta of the manifest to request each changed file, obtaining the appropriate deltas to get up to current.
Any holes in this scheme? Any suggestions for better options?
In a nutshell, I've got over 100,000 individual files that need to be tracked and synchronized across several hundred different client locations, and potentially many more clients than that in the future. Old versions of the files need to be available so that outdated clients can patch their way up to current; this is already supported via a delta system but needs to be integrated into a larger architecture. Essentially, I need to generate a manifest of all the files, update this manifest live when files change on an authoritative server, and distribute the manifest (and accompanying files) on demand to clients as they patch up to current.
My overall plan involves something like this:
- Generate a CRC32 of each filename and use this as a hash index for fast lookups of files
- Fall back to canonical filename if the CRC32 hits a collision
- Serve files and the manifest using a custom HTTP server with a virtual file mapping from URIs to the file system/manifest metadata
- Allow serving patches of the manifest given a manifest version ID and a target version ID
This should permit the manifest to grow very large (i.e. hold metadata for all of our files) without requiring that it be distributed in its entirety every time a client needs to patch. A client simply first patches its manifest, then uses the delta of the manifest to request each changed file, obtaining the appropriate deltas to get up to current.
Any holes in this scheme? Any suggestions for better options?