Original Post
Relational Database + Versioning Filesystem
The hierarchical filesystem model that is so familiar to so many of us is tired and inadequate. It is a useful tool for grouping data by function or property, but being constrained to place all of our data into strict single-habitation hierarchies limits our ability to manipulate our data in powerful ways.
Data Organization: Relational Database
Say, for a moment, we completely discarded the hierarchy and stored all of our files - applications, scripts, images, audio, video, text - in a linear or flat repository. That''s not really useful because it becomes hard to find anything or work on groups of data. Say we then allowed the user to define relationships or associations between data/files and abstract concepts. Say we allowed the user to define multiple such associations between a datum and different concepts. Now the user can use the associations as pivots for viewing data:
But that''s not all.
Version History
Journalling filesystems are widely available today (XFS, reiserfs, JFS, NTFS). What we are proposing here is a filesystem that stores each file as a base image and a sequence of deltas, effectively providing rollback capability. Every so many deltas, a new base image is inserted and deltas are created from there. The user may specify that a previous image and its sequence of deltas be discarded if older than a certain data, or if filesize exceeds a certain amount, or after a certain number of new base images, or after a given period of inactivity, etc. Think a mutated CVS built into the filesystem (or ARMS; in-joke for Dean Harding).
Version history is closely tied in with the expanded concept of permissions that RDVfs (working title) would implement. Rather than Unix-style 10-bit permissions, file accesses would be specified in terms of roles - author, co-author, consumer, maintainer/admin, etc. Multiple group and user permissions would be possible, allowing different users to have different levels of access independently of any groups, and all permissions are being designed with a view to maximizing the ability of users to publish and share their data. Time-variant permissions would also be possible, allowing a user to specify temporary permissions or permission changes for other users and groups.
The intent of this post is to canvas for opinions, suggestions, criticisms and critiques. Hopefully it will be the first in a series of threads on ways to revolutionize modern computing (ambitiously pseudo-titled "Next-Generation Computing"). Thank you for your time.
Our user is doing a media research project on character realization in modern media - film, roleplaying games and books. The user creates an association labelled "Characters" for all the files describing various figures from diverse media. The user then creates another association, "Film" for characters from film media and similar associations for "RPG" and "Print". Our user can now choose to view all characters (the filesystem is being searched with "Characters" as association criteria), and receives a list of all characters as well as subcategorization of some entries as pertaining to roleplaying games, film or books. The user then selects both RPG and Film and views the results (the filesystem is being searched with "Characters", "RPG" and "Film" as association criteria).The idea isn''t terribly revolutionary; I believe Windows Longhorn is to feature a filesystem something like this. Th total abolishment of the hierarchical filesystem (relegated only to being a viewing abstraction) has powerful implications for file location, though, as a file is now never more than two levels "away" from the current location. This dispenses of the need for filepickers, for one thing, and makes specialized search tools less necessary. File associations (analogous to file locations in hierarchical filesystems) become metadata along with file type, size and modification data. Provided is an example implementation (C++) of rudimentary functionality and a brief (and minimal) test application exhibiting the association and criteria pivoting functionality:
// repository.h
#include <list>
#include <string>
#include <vector>
typedef size_t GUID;
class Repository
{
private:
struct entry
{
GUID inode;
std::string identifier;
std::list<GUID> relationships;
//*** TODO: add storage capacity (reference)
};
typedef std::list<entry> repository;
typedef std::list<GUID> relationship;
repository contents; // files
repository relationships; // folders
struct id_match
{
GUID comp_id;
id_match( const GUID id ) : comp_id( id )
{
}
bool operator () ( const entry & e )
{
return (e.inode == comp_id);
}
bool operator () ( const GUID g )
{
return (g == comp_id);
}
};
public:
Repository() {};
~Repository() {};
void addFileRelationship( const GUID file, const GUID folder = 0 );
void addFolderRelationship( const GUID subcategory, const GUID category );
void addFolderContentsRelationship( const GUID from_folder, const GUID to_folder );
GUID newFile( const std::string & file );
GUID newRelationship( const std::string & folder );
void setFileIdentifier( const GUID file, const std::string & identifier );
void setFolderIdentifier( const GUID folder, const std::string & identifier );
std::list<GUID> getFiles( GUID in_folder );
std::list<GUID> getFiles( std::vector<GUID> in_folders );
std::list<GUID> getFolders( GUID in_folder );
std::list<GUID> getFolders( std::vector<GUID> in_folders );
const std::string & getFileIdentifier( GUID file ) const;
const std::string & getFolderIdentifier( GUID folder ) const;
static GUID getNextFileID( void );
static GUID getNextFolderID( void );
};
// repository.cpp
#include "repository.h"
#include <algorithm>
void Repository::addFileRelationship( const GUID file, const GUID folder )
{
repository::iterator file_it = std::find_if( contents.begin(), contents.end(), id_match(file) );
repository::iterator folder_it = std::find_if( relationships.begin(), relationships.end(), id_match(folder) );
if( file_it != contents.end() && folder_it != relationships.end() )
file_it->relationships.push_back( folder );
}
void Repository::addFolderRelationship( const GUID subcategory, const GUID category )
{
repository::iterator subcat = std::find_if( relationships.begin(), relationships.end(), id_match(subcategory) );
repository::iterator cat = std::find_if( relationships.begin(), relationships.end(), id_match(category) );
if( subcat != relationships.end() && cat != relationships.end() )
subcat->relationships.push_back( category );
}
void Repository::addFolderContentsRelationship( const GUID from_folder, const GUID to_folder )
{
repository::iterator file_it = contents.begin(), file_stop = contents.end();
while( file_it != file_stop )
{
relationship::iterator from_it = std::find_if( file_it->relationships.begin(), file_it->relationships.end(), id_match(from_folder) );
relationship::iterator to_it = std::find_if( file_it->relationships.begin(), file_it->relationships.end(), id_match(to_folder) );
if( from_it != file_it->relationships.end() && to_it == file_it->relationships.end() )
file_it->relationships.push_back( to_folder );
++file_it;
}
}
GUID Repository::newFile( const std::string & file )
{
entry e;
e.identifier = file;
e.inode = getNextFileID();
contents.push_back( e );
return e.inode;
}
GUID Repository::newRelationship( const std::string & folder )
{
entry e;
e.identifier = folder;
e.inode = getNextFolderID();
relationships.push_back( e );
return e.inode;
}
void Repository::setFileIdentifier( const GUID file, const std::string & identifier )
{
repository::iterator it = std::find_if( contents.begin(), contents.end(), id_match(file) );
if( it != contents.end() )
it->identifier = identifier;
}
void Repository::setFolderIdentifier( const GUID folder, const std::string & identifier )
{
repository::iterator it = std::find_if( relationships.begin(), relationships.end(), id_match(folder) );
if( it != relationships.end() )
it->identifier = identifier;
}
std::list<GUID> Repository::getFiles( GUID in_folder )
{
std::list<GUID> ret;
repository::iterator file_it = contents.begin(), file_stop = contents.end();
relationship::iterator it;
while( file_it != file_stop )
{
it = std::find_if( file_it->relationships.begin(), file_it->relationships.end(), id_match(in_folder) );
if( it != file_it->relationships.end() )
ret.push_back( file_it->inode );
++file_it;
}
return ret;
}
std::list<GUID> Repository::getFiles( std::vector<GUID> in_folders )
{
std::list<GUID> ret;
int N = in_folders.size();
for( int i = 0; i < N; ++i )
{
std::list<GUID> r = getFiles( in_folders[i] );
ret.splice( ret.end(), r );
}
return ret;
}
std::list<GUID> Repository::getFolders( GUID in_folder )
{
std::list<GUID> ret;
repository::iterator folder_it = relationships.begin(), folder_stop = relationships.end();
relationship::iterator it;
while( folder_it != folder_stop )
{
it = std::find_if( folder_it->relationships.begin(), folder_it->relationships.end(), id_match(in_folder) );
if( it != folder_it->relationships.end() )
ret.push_back( folder_it->inode );
++folder_it;
}
return ret;
}
std::list<GUID> Repository::getFolders( std::vector<GUID> in_folders )
{
std::list<GUID> ret;
int N = in_folders.size();
for( int i = 0; i < N; ++i )
{
std::list<GUID> r = getFolders( in_folders[i] );
ret.splice( ret.end(), r );
}
return ret;
}
GUID Repository::getNextFileID( void )
{
static GUID last_file_inode = 1; //*** TODO: replace with GUIDs
return last_file_inode++;
}
GUID Repository::getNextFolderID( void )
{
static GUID last_folder_inode = 1; //*** TODO: replace with GUIDs
return last_folder_inode++;
}
// main.cpp
#include <iostream>
#include <iterator>
#include "repository.h"
int main()
{
using namespace std;
Repository r;
GUID apple = r.newFile( "apple" );
GUID bicycle = r.newFile( "bicycle" );
GUID car = r.newFile( "car" );
GUID cat = r.newFile( "cat" );
GUID dog = r.newFile( "dog" );
GUID eggs = r.newFile( "eggs" );
GUID orange = r.newFile( "orange" );
GUID plane = r.newFile( "plane" );
GUID train = r.newFile( "train" );
GUID Animals = r.newRelationship( "Animals" );
GUID Food = r.newRelationship( "Food" );
GUID Fruits = r.newRelationship( "Fruits" );
GUID Vehicle = r.newRelationship( "Vehicle" );
GUID Wheeled = r.newRelationship( "Wheeled Vehicle" );
GUID Winged = r.newRelationship( "Winged Vehicle" );
r.addFileRelationship( apple, Fruits );
r.addFileRelationship( bicycle, Wheeled );
r.addFileRelationship( car, Wheeled );
r.addFileRelationship( cat, Animals );
r.addFileRelationship( dog, Animals );
r.addFileRelationship( eggs, Food );
r.addFileRelationship( orange, Fruits );
r.addFileRelationship( plane, Wheeled );
r.addFileRelationship( plane, Winged );
r.addFileRelationship( train, Wheeled );
r.addFolderContentsRelationship( Fruits, Food );
r.addFolderContentsRelationship( Wheeled, Vehicle );
r.addFolderContentsRelationship( Winged, Vehicle );
std::list<GUID> l = r.getFiles( Vehicle );
std::copy( l.begin(), l.end(), ostream_iterator<GUID>(cout, "\n") );
l = r.getFiles( Food );
std::copy( l.begin(), l.end(), ostream_iterator<GUID>(cout, "\n") );
return 0;
}