Wednesday, January 21, 2015

23 MS-Windows ® based cataloguing – HomeBase ® at the home base!

It would be nice if they could develop a basic program that has Lotus Agenda’s unique approach and capability (see previous post) of assigning items to categories based on text matching, as this is what automated classification is about. Until then, one has to fall back upon other database programs available for inventory management. One such, especially suitable for home libraries, is a simple and effective program called HomeBase, which is available free from AbeBooks at abebooks.com (American Book Exchange, if I am not mistaken). It is actually meant to make your book catalogue available to prospective customers on the AbeBooks site (for a payment); but we can use it in the meanwhile to develop our own stand-alone catalogues. In doing so, we may have to find certain work-arounds to emulate Agenda’s sorely-missed capabilities.

HomeBase (now in version 3) has fields for most of the details one would like to enter for a book, and then some. Being a database for book sellers and collectors, it has fields for binding, type of edition, condition of the book, size, and so on. Many of the fields are sortable, in the sense that the display or list View can be rearranged alphabetically in ascending or descending order according to Author’s name, or Title, or Publisher, to name a few possibilities. It also has fields for keywords, comment/description, and private notes, ISBN, etc., but since it’s obviously not geared to the classified library, it does not have a field specifically for the Dewey classification code or call number /shelf number. An obvious choice to enter the DDC number would be the Keywords field, and that is what I do, but it’s unfortunately not one of the sortable fields. Since the catalogue needs to be arranged by the DDC number (apart from the Author name), I suggest using one of the available fields for this; I am presently experimenting with Illustrator, which is a sortable field: you can display the list in the order of the DC numbers on this field with a key stroke; the Description/Comment and Keyword fields do not have this capability, which makes them less useful if you wish to arrange the books by Subject or Class Number. If there are a lot of books under each  subject or class number, it would obviously be good to put the full Shelf No in the DDC field (DDC number, three letters from Author name, year), so that the books can be arranged in the same specific order on the shelf as well as in the database view.

The plus point is that books can be picked up based on text searches in the Keyword and Description/Comments fields (and also based on Book Number, Title, Author/Illustrator, Publisher, ISBN, and Status fields); this means that you can get the program to list the books having a certain text string in any of these fields (termed a Filter), say ‘Wildlife’, and then sort them by Author or Year and so on. Incidentally, there is also a long list of category names already built in, so you could use these instead of the Dewey subject headings. Or you could add your own categories (up to a limit, I think it is 100).

All this is not as versatile as Lotus Agenda, where you could type any text desired  in the main Item field, and Agenda would automatically make assignments of the Item to different (pre-existing) categories based on text matches. In HomeBase, I don’t think there is a facility to save specific queries as Views (which is another of Agenda’s delightful features), one has to ask it to pick out items matching your criteria, and then work from there. But there is another facility in HomeBase that should be useful: you can assign each item to different Catalogues. One use may be to put in DDC Classes here; there is of course a limit to the number of catalogues, something like 100 I think. The idea would be, I suppose, to have a limited number of Catalogue names that could broadly follow the DDC Hundreds (with a few selected sub-disciplines to reflect the local interests, if there were a large number of specialized books). If the file starts getting too big, for instance, one may think of hiving off portions of it, say all Humanities (000-499) in one Catalogue, Science & Technology in another (500-699), and so on.

Another potentially useful feature in HomeBase is that it will locate the book in its own database and fill in all the fields if you give it some information like the international standard book number (ISBN). That requires you to register as a seller, however, which starts at 25 $ a month for 500 books (apart from a commission on sales), so you had better be a serious vendor with a good inventory and be prepared to work hard in case you want to recover your money! I guess eBay is a little easier to start as a small or occasional seller, as it allows you to list up to 50 items free and charges a 10% commission only on sales.


Disclosure: I’ve only played around with HomeBase, and not actually gotten down to filling in the database records for either my books or my music albums. I think that necessity will arise only if I intend to sell (and that too, through AbeBooks). I’m not very sure that it will be worth the effort at present to fill in all the books, merely to be able to search and locate automatically… my collection is not that big that books get completely lost sight of if they are misplaced on the shelves!

Monday, January 12, 2015

22 Building catalogues with a Personal Information Manager (PIM) - Lotus Agenda ® for MS-DOS ® !

Have computer, will database. Obviously, a book catalogue is a prime candidate for computerisation. Funnily enough, one of the greatest software programs I have ever used for this type of application, is an old MS-DOS package by Lotus Corporation called Lotus Agenda. Here’s what Underdogs has to say about this program on the site http://www.hotud.org/component/content/article/40-application/23496

“Lotus Agenda is arguably not only the best "personal information manager" (PIM) software ever made, but also one of the best applications ever seen on a PC. A DOS program originally marketed by Lotus during the late 1980's and early 1990's, Lotus Agenda is in fact the program for which the term PIM was coined. Even to this day, it has PIM capabilities and features that are unmatched by any other software available.

"Until recently Agenda was the only PIM in the market that allows the keying of data to precede the creation of database tables. It is an immensely useful tool for sorting piles of information into meaningful categories.” 



The great thing about Lotus Agenda is the capability of searching out matches and making allotments to categories automatically based on these text matches. In specific terms, this can be used to assign each book to its subject classifications (which can be more than one), based on text matches – words or parts of words. All other programs I have tried require you to make assignments manually – this job may be marginally speeded up by drop down lists of categories or by autocompletion of the entry based on your initial key strokes, but you would still have to look at the list of categories (say, Dewey numbers) and choose the appropriate entries manually. What Lotus Agenda does is to parse through the description (your main item entry), and match it with the words in its list of categories , and automatically make assignments (which you could overrule manually if required). Another, probably unappreciated, and unexpected, advantage of this is that each item can have a different number of assignments. In a usual database program, it is generally required to specify a fixed set of entries under each category. For instance, we would need to specify Author1, Author2, and Author3, in the design stage of the database. Similarly, Subject1, Subject2, Subject3 and so on, but a definite fixed number. Now if a particular item had say four authors, we would be stuck. These traditional databases are fixed in their structure. Lotus Agenda, on the other hand, can have a different number of assignments under each head, for each item.

Of course, we would have to provide the master list of categories (and sub-categories too, nested a few levels). Each file can get to 5.5 megabytes (MB) according to the Help pages, and you would have to split the file beyond this size. I used it for a collection of 2000 items with close on a 100 categories (average 25 categories assigned to each item!) and the file size hardly came to 2 MB. This is suitable for the simplified Dewey classification, but because very big file sizes generally increase errors in processing (the file will come out “corrupted” sometimes, and we will have to reload the older file), I have generally split up the catalogue by broad areas of knowledge, e.g. each 100’s gets its own database file. 

The advantage of multiple assignments is that a book can be classified in many ways. A book on wildlife conservation could thus be automatically classified under Wildlife management in 639.9, and simultaneously also under other subject categories (DDC subject codes) like Conservation of Biological Resources 333.9, Animals/Zoology  591, Ecology 574.5, etc. The pre-condition for this to happen automatically would be either that we enter the numbers in the description itself, or that the category names each should have the keywords of the item description (Title of the book, etc.) already entered in its definition. The flexibility afforded is that keywords can be added to the definition at any time if some new text items crop up as the books are entered. . It is doubtful whether a single file could handle the entire DDC subject categories as well as the Tables for Place, etc. I guess what one should be doing is to add subject headings and place names as one goes along, which will help to keep file sizes down, and not really try to pre-enter the entire DC subject lists.

The disadvantage – and a crushing one, I’m afraid – is that Agenda was only made for DOS, and never ported to Windows. There are groups that discuss these things on the Internet, and they used to suggest that an Open Source project was under way led by Mitch Kapor, one of the original team that developed Agenda for Lotus. Unfortunately it appears that they were interested in certain other capabilities, especially web-based information tracking and team-based processing (emails, contacts, calendars  and the like), and it seemed to have resulted in ‘bloatware’ that doesn’t do the intelligent Personal Information Management (PIM) that a single user needs.

One good thing that Lotus has done is to provide much of the old software as freeware. The official download page for Agenda (from the above blog, dated March 2009) is http://www2.support.lotus.com/ftp/pub/desktop/Agenda/dos/2.0/misc/
The site provides the individual installation disks as a file image, a throwback to the 3.5 inch microdiskettes on which they used to come (each had a 1.4 MB capacity!).

One really wishes there were an alternative ‘free-form’ database package where one could change the structure, add fields, have variable number of fields for each item (record), and have the software intelligently parse the text and make matches on the run, but there doesn’t seem to be any for Windows. A German product called InfoHandler seems to have made the effort, but I could never understand the logic of their system and gave up. The one practical option seems to me to be ABE (American Book Exchange) HomeBase, which they provide free of charge here http://www.abebooks.com/homebase/software-inventory-management-system-catalog/?cm_sp=Ftr-_-Home-_-C2

I preferred v.2.x, as v.3 which they have now has very tiny typefaces! One advantage if you pay the membership is that the program can fill all the fields from its international database if you give the basic information. You could also be a seller on AbeBooks.com if you pay the fees. More on AbeBooks HomeBase in the next post! 

Abundant disclosure: I used Agenda mainly for my music collection, and only experimented with it for the Dewey classification. Which eminds me, I really must port the music data into HomeBase... but I will have to make assignments to the categories manually!

Monday, December 29, 2014

21 Dewey on the wild side

This one is on classifying Wildlife and related topics in the Dewey Decimal system. Just as with Forests (see posts 12-14), Wildlife also poses the problem of too many choices! And these numbers are situated variously in the Social Sciences (under 333, Land and natural resources), in 639 (Hunting, fishing, conservation, and related technologies), and under various numbers in the Biology sections. Let’s have a closer look.

Say you just got  a copy of a lovely book on the Wildlife of the Indian Subcontinent, about all the richest wildlife habitats and the habits and conservation status of the important animals and birds, their place in history, religion, and culture, and so on. Where would I like to put books on the wildlife of this place or that on my shelves? My first instinct would be to… follow my instinct! I think it would be my instinct to gravitate to the biology shelves… but here we have a problem, because the book can be filed in Animals (590), or in Ecology (577), or in Natural history of organisms (578). The biology numbers like 590 may feel a bit hard-core (in the sense that they are for more scientific or zoological treatises on body parts, for example), whereas we are looking for a place to put works for the animal-lover and watcher of live animals (often the opposite of the biologist!). This is what is called “natural history”, not quite official as far as the hard-core are concerned, but DDC 22 has fortunately provided a nice alternative in the form of 578, Natural history of organisms (which is a relief from DDC 19 which sent you to 508 for Natural history). The strange problem here is that they don’t seem to provide for geographical faceting under this particular number. They prescribe 578.01-578.08 for “standard subdivisions”, then provide only 578.09 for “Historic, geographic, persons treatment”, but don’t mention extensions of -09 for specific locations and jurisdictions (578.093-099, as they usually do in their schedules), but only show one entry, 578.0999 for “Extraterrestrial worlds”! They do have a caution not to use 578.0914 to 578.0919 extensions for general regions, but instead to use 578.73 to 578.77, under which you have various ecological types like forest, grassland, etc. (repeated from 577.3-577.7, under Ecology). I pay no heed to this implied truncation of -09 numbering, and go right ahead and form the numbers like 578.0954 (for the Indian sub-continent, for example). And, naturally, other similar numbers for all “Wildlife of…” type of books which deal with all types of animals and birds, in relation to the climates, habitats etc. of regions and countries in general.

The section 578 also has special subdivisions for other types of natural formations, under 578.7, “Organisms characteristic of specific kinds of environment”: the numbers after 577 from Ecology, 577.3 to 577.7, are added to 578.7. Forests, for instance is 578.73 (from 577.3, Forest ecology). So books on “Rain forests” will go in under 578.734 (from 577.34 Rain forest ecology), and you can always append geographical endings using 09 from standard subdivisions.

The matter doesn’t end there, however (how could it be so straightforward!), as you may like to use numbers under Animals (590) or Mammals (599) or Botany 580 or whatever, for specific “taxonomic groups”. Say you have a book on the “Large mammals of Africa”, that is rhino and elephant and lion and so on: would you like to put it under 578, or would you shift it to its own niche in 599.1, Natural history of animals? Similarly for other groups. You may like to put a book dealing with the botanical aspects of forests under 581.73 (again this repeats the numbers from 577.3 to 577.7), rather than under general natural history. That is, you could choose to differentiate the books depending on their focus, or accent: is it dealing with all sorts of organisms? Does it describe the whole ecosystem or does it talk of each species in particular? The latter would be better off in the narrower number referring to the taxonomic grouping: say, a “Field guide to the mammals of India” would go under 599, rather than 590 or 578, which could be for books dealing with their ecological relationships.

Another type or genre is books on behaviour, ethology. Previously Ecology and Ethology used to be treated pretty closely together. Now the choice would be to put Behaviour under the specific sub-class under the taxonomic group: “Behaviour of mammals” under 599.15, of Birds under 598.15, of Animals under 591.5. You have sub-divisions under them for different aspects of behaviour, such as territory, feeding, mating, nesting, migrating, and so on (they have omitted 581.5 for Behaviour of Plants, presumably expecting us to be happy with 581.7 Plant ecology).

As if this weren’t enough, you have a totally different set-up under Technology, 639.9 (Conservation of biological resources), which comes after Agriculture, Horticulture, Forestry, Animal culture, and so on. This suggests a differentiation of techniques of husbandry from basic knowledge of the organisms. Now is wildlife management a form of husbandry or a subset of ecology? Under 639.9, they have headings like 639.92 Habitat improvement, 639.93 Population control, 639.95 Maintenance of reserves and refuges, 639.96 Control of diseases etc., 639.97 Specific kinds of animals, and so on up to 639.979 for Mammals and 639.99 Conservation of plants, which suggests what types of topics go here. I tend to file the more technical books and reports on wildlife here: manuals on census operations, manipulation of habitat, captive breeding, disease management, policing (a part of protection), plans and reports on wildlife parks and congresses, and so on. There is a category of books which I am still vacillating about, puttng them at times under 578, at other times under 639.9: this is books on specific wildlife parks and sanctuaries. The profusely illustrated series of collector’s volumes published by Sanctuary magazine, for instance, on individual wildlife areas (Corbett, Bandhavgarh, Sunderbans, and so on), and some imitators, for instance, treat of the wildlife of the region and should go under 578, but I prefer to have them under 639.95, Wildlife reserves, because they are actually focused on the management of these particular jurisdictions, each with a unique background, history, and set of problems and solutions. I feel these are books primarily useful for the wildlife manager (639.9), although packaged as a table-top picture book for the general wildlife enthusiast (578). I guess either choice would be acceptable. General accounts of wildlife parks (protected areas) in a state or region also go under 639.95, even though they may describe their habitats, give species lists and talk about the habits and ecology of the organisms.


We’re not done yet: there is still the disturbing factor of the social sciences, which we met with 333.75 Forests, and now meet again under 333.95 Biological resources (conservation of). Many CIP (Cataloguing-In-Publications) entries I have noticed, tend to put all multi-disciplinary accounts under 333 (Economics of land and energy) sub-divisions, as recommended by Dewey: especially the types of books published by National Geographic. I tend to avoid this, unless we are dealing specifically with the social or economic aspects. A book on Wildlife economics, for instance, or books dealing with wildlife and tribal rights, or community management, or international conventions, or policy, may prefer this location. On the other hand, there is a tendency to send Nat Geo books equally to Geography & Travels 910 to 919, or Ethnology or Human ecology (indigenous peoples and so on) to 306. There could be other subdivisions on specific aspects like Government and Public administration, Law, International cooperation, Trade, Commerce, Production, Non-governmental or Voluntary organizations, etc., which may receive some of the books and reports, especially boring annual reports and ministry documents. In all this, finally, we may have to choose two (or at the most three) favoured locations, even if there were other tailor-made choices, in the interests of keeping stuff together on the shelves.

Saturday, December 13, 2014

20 The physical catalogue on cards

We would all like to have a detailed list of the books and tapes we own. The basic version, of course, is to have a long notebook (a ledger) for each type of possession, and go on entering our acquisitions as they come in, with basic description, title, date of purchase, and price paid. A separate ledger could be maintained for books and other texts, perhaps with separate sections for periodicals and for reprints or ‘grey’ matter (newsletters, mimeographs, occasional documents); and separate ledgers for recorded media (CDs, tapes, etc.), all types of equipment, and what have you. A running serial number may be all that is required to identify each item, and if you put this number on a sticker on or in the item itself, you have a robust and simple system to keep track of their status. When you give an item away, you can record the information and draw a diagonal line through its entry as token of disposal. In fact I use precisely this system to keep track of my financial investments (and significant equipment purchases), as it has the advantage over a computer based system of being always ready to go, robust and physically available at hand, and amenable to all sorts of annotation, on the run, whenever a thought strikes. Of course, it doesn’t produce nicely formatted reports or column totals, and doesn’t send out warning beeps when it’s time for renewal or servicing, which a computer system could do, but I suspect that it will be too late by the time I get round to putting all this on disk. Anyway, the old data will always reside between the covers.

When it comes to books, however, if you plan to have a few thousand, it makes sense to build up a card catalogue from the beginning. My card catalogue started when I was collecting references for my doctoral thesis; my book acquisitions took up steam only sometime after that, so it was a natural extension to enter the books as well on those 3 by 5 inch cards. Now it has become a ritual whenever I get home with any books or reports, whether from the bookshops or from meetings and conferences. They all get entered in the 3 by 5’s, and put into the card tray. Usually I enter the classification number as well, but if I am too bothered with other stuff to do it rightaway, I keep the unclassified cards in a separate holding tray, to be filled in later. I also enter the classification and date of purchase on the first leaf of the book itself (in pencil!), so that I can put it in its due place on the shelf after I have finished reading it (which is falling behind these days!.

What do I put on the card? There are very elaborate conventions on this, the best known being the Anglo-American Cataloguing Rules (AACR2), of which I have a copy of the Concise version, revised 1988, prepared by Michael Gorman, and published jointly by the American Library Association (Chicago), the Canadian Library Association (Ottawa), and The Library Association (London). But I rarely look into it. I have standardized on the following format: leave the top line blank, on the next line enter the author’s name following the usual last name – first name conventions used in citing references, and year of publication; below that, the title of the work and any subtitles or smart one-liners; then other editorial information like series or set name and general editor (if important enough), illustrators, foreword writer (if an eminent person), then edition number, publisher and place, and finally the international book number (and Library of Congress number if available). At the top left, I write the Dewey class number, followed by shelf numbers (usually three letters from the author name, followed by year) to identify it uniquely; on top right, any special Location (Music Records, or Series, or Loft, for example). At bottom left, I pencil in date and price (both original and buying price if needed), and any supplemental information like key words, alternate classification numbers, etc. All this by hand: it takes a couple of minutes, and my record is ready! The cards are physically kept in a metal card cabinet with four sliding trays. You don't even have to go and buy the printed cards: you could do as well with any old paper cut to size (I notice my institute library does this for their internal purposes, although their actual catalogue is on computer, of course). Not pretty, but works well enough!

The official AARC rules are very precise about what each card should contain, and they also have official registers for the correct way of expressing names and so on. For the record, the following areas are prescribed:

Area 1: Title and statement of responsibility
Area 2: Edition
Area 3: Material (or type of publication) specific details (serials, computer files, maps, music etc.)
Area 4: Publication, distribution, etc.
Area 5: Physical description
Area 6: Series
Area 7: Notes
Area 8: Standard number and terms of availability
Area 9: Supplementary items
Area 10: Items made up of more than one type of material
Area 11: Facsimiles, photocopies, other reproductions

One point on which I disagree with the AARC2 is the rule that editors and compilers should not be made the “main entry”. I prefer to stick to only one type of main entry, which is the author or editor, and if this is not available, then sometimes the corporate body itself or even the publisher (like Government, or National Geographic, or Newsweek, or Oxford). AARC2 says that in the absence of a clear author or creator, one should use the title as first entry (leaving out articles at the start, e.g. Oxford Dictionary of Quotations, The). Sometimes I have used the dreaded Anonymous, too, but that is not a happy solution as it may tend to bunch up a lot of stuff at the head; much better to put the organisation name instead.

Actually, I don’t think they expect all the fields to be filled; the main bits, of course, are author, title and identification by edition or book number. If you have the time, by all means fill in some of the other stuff. The class numbers are my favourite, because I arrange both my cards and my shelves according to them; naturally, I favour Dewey Decimal numbers (I’m on DC22 now). This gets the books in order of field of knowledge and subject matter (in the Dewey order with all its idiosyncrasies!), which suits a knowledge-based user better than arranging by author name alone (or by title!). Very occasionally, if a book seems equally at home in two classes, I may put in a card for each DDC number, giving the shelf position on the top line. Of course, if I have two copies (which happens occasionally!), I put one copy in each location.

These other types of catalogue, of course, are also useful sometimes (e.g., if you are making up a short list in a particular discipline). In public libraries, they used to make up two card catalogues, one arranged by Subject (following the DDC or any other system), and the other by Author, called respectively the Subject Index (which could be an Alphabetical or a Classified Index) and the Author Index. Nowadays, of course, catalogues are maintained on computers, and the database software will allow you to list them by almost any of the fields: maybe by year, or publisher, or combinations.


I don’t actually use the card catalogue much, except to keep it up to date. At the back of my mind is the expectation that I will enter it into a computer some day (but I wonder whether that will actually be useful). I do not think it will help my heirs to sort out what is to be thrown or given away, nor do I expect my Maker to call me to account on this matter! I do riffle through it once in a while to see whether I have a certain book already, if I cannot see it anywhere around. Of course, since it is only a classified index, it won’t help me do alternate searches on author or ISBN or title; that will be possible only if it is put on a computer. I did make an experiment with a couple of software packages to do this, and I will talk about this next post.

Tuesday, December 2, 2014

19 Classify and Catalogue – two sides of a page

One small point that may be worth pointing out here is the distinction between two parts of the process: one is cataloguing, the other is classifying.

A catalogue (catalog) is basically a list of all items: they could be our own possessions, or they could be even a seller’s stock list (or inventory) or our own wish list. When it comes to books, there are usually two basic types of lists used: one is an Author Catalogue, which starts each item with the author’s name, followed by whatever details we feel are needed to identify the item uniquely. The other type of list is a Subject catalogue, where the first entry is the code number or word for the subject. If anybody remembers, our institutional libraries used to have these two types of catalogues written on cards of approximately 3 by 5 inches, stored in sliding trays that formed an impressive piece of furniture near the entrance. Earnest scholars would spend hours thumbing through these cards, trying to locate their particular requirements.

Both types of catalogues have their uses. If we wish to locate records (cards) for books by a particular author, say Dickens, and we do not know where in the library shelves these books will be found, we go the Author Catalogue and pull out the tray for the D’s. Of course, the utility of these cards depends on what else is recorded on each book’s record: usually it includes the serial number of the book in the library’s stock register (the Accession number) which should be a unique identifier, the book title, year of publication, and publisher’s name and edition, and perhaps the international standard book number (ISBN). There are detailed codes written for this (the Anglo-American Cataloguing Rules AACR2, for instance) But apart from these, the most useful to locate the book would be a location number or code. In the least ordered library, they could be simply stacked in the order of their accession numbers, as we suggested could be done for reprints of papers, as they don’t have solid spines which will enable them to stand up on their own in the shelves. For most collections, however, it will be nice to have them grouped by subject, which is where the Classification scheme comes in.

Commercial bookshops usually do not adhere to a very strict code of classification, and generally group books by broad subjects, like Physics, Chemistry, Sociology, Politics, History, Current Affairs, and so on, maybe even under further subdivisions if it is a campus bookshop, say different branches of Chemistry or Physics or whatever (usually following the university syllabus). Meticulous (let’s face it: somewhat obsessive-compulsive) documentation experts like Dewey (or to a greater degree, Ranganathan of the Colon Classification) like to reduce these non-standard subject headings to a standard code with a consistent system of labeling. Dewey, of course, uses mainly numbers: the 3-digit numbers stand for the one thousand subject headings (Sections), which are then expanded by adding on further digits to the right after a decimal point (the successive digits show the hierarchical position, rather than a value). The Colon Classification system follows a different philosophy, which I won’t even try presenting here (maybe another day!). It is my impression that the Dewey Decimal system is more popular because it has a strong management backup with the OCLC Online Computer Library Centre, Inc. (obviously), is constantly being developed by specialists at the Library of Congress, and more than anything has an infinitely more friendly and common-sense approach as against the Colon’s (let us face it) somewhat dry language, exaggeratedly punctilious rules and cryptic terminology (well, one has to admit that it lives up to its rather unfortunate name).


The Dewey number locates the book under the broad discipline (more or less the 3-digit Section heads, but also maybe under further subdivisions where there are distinct areas), and then under the most appropriate specific subject according to the entries in the schedules. However, for large collections, one can go even further, by adding various suffixes from the half-dozen Tables provided in Volume I. These enable various aspects or facets to be specified: a favourite is the geographical or regional coverage, for instance, from Table 2. The Standard Subdivisions in Table 1 provide tags for various aspects: the suffix -01, “Philosophy and theory”, for example, is the first subdivision provided in many numbers in the schedules (it could be used for Policy as well). The suffix -09 introduces the geographical locations, persons, and periods from Table 2, and so on. Table 3A and 3B provide tags for different numbers of authors and so on (for collections and anthologies, for instance). Table 3C provides “Additional Notation for Arts and Literature” to be added “where instructed”. Table 4 provides “Subdivisions of Individual Languages and Language Families” (from 400), Table 5 “Ethnic and National Groups”, Table 6  provides tags for “Languages”, while Table 7 for types of persons has been deleted.

Friday, November 28, 2014

18 Why classify when “Everything is miscellaneous”?

While on this topic of classifying our music resources, it would be as well to revisit the question of why at all we want to have a classified list of our possessions. Why not just keep a running list where we enter each thing as it comes, something like a “general ledger account” of day-to-day transactions?

I came upon a very interesting book on this question (yes, there are geeks who write whole books on as mundane an activity as classifying and arranging!) titled “Everything is Miscellaneous – The Power of the New Digital Disorder”, by David Weinberger (published 2007 in Times Books by Henry Holt & Company, New York, ISBN 978-0-8050-8043-0). Weinberger is described as a fellow of the Harvard Law School’s Berkman Center for the Internet & Society, and an adviser and consultant for Fortune 500 companies, bestselling author, and a doctor of philosophy, so he should know a thing or two about the subject. His thesis is that the power of the computer and the Internet have placed huge databases at the call of a button, and searches on key words can be made in a fraction of  a second, so nothing really needs to be classified any more. In other words, everything can be entered in a single, massive list of all things.

The prime example Weinberger cites is, appositely, from the world of online music resources, the Apple iTunes music store, where the albums I referred to in the previous post, have been unpacked as tracks, enabling consumers to download just what they want, when they want, thereby giving Apple “more than 70 percent of the market”. The files need not be organized on your hard disk in any sort of structure, as the software allows you to search on key words any time (it will be faster, I suppose, if all the filenames and their embedded metadata (data about data!) are already indexed on every word, like Google does for the entire Internet). By giving keyword filters, we can narrow down the choice as much as we wish. Even filenames need not be intelligible (say Western-Bach-Concerto-Violin-No.1), but can just be a miscellaneous number, as the file’s embedded metadata would have information on which searches and selections can be made.

Another example Weinberger likes is the way digital cameras have generated millions and millions of pictures that are stored on the Web, viewed and exchanged and commented upon on the Web, and seldom printed out or put into albums. Here again, programs like Flickr.com have relieved the consumer or average user of the responsibility of classifying the pictures – as long as some metadata is embedded to indicate date, or maybe location, even the subject may not be all that relevant. What Weinberger says is that in order to “take full advantage of the digital opportunity – we have to get rid of the idea that there is a best way of organizing the world”.

On the other hand, if you tried this with physical objects – your collection of photographic prints or slides, or your books, or your CDs – you would end up with a big pile of stuff that you would have no way of using purposefully unless you did some sort of sorting – usually by subject, not colour or size! There actually was such a bookshop in Bangalore, where I live, which was more or less an icon – the owner could locate almost any book in his pile, but few others could. The store, sadly, closed a few years back, but another phenomenon has sprung up in many cities across the country – shops selling huge collections of used books imported by the carton (we are talking of shipping, not cardboard), sometimes even sold by weight! Most of these stores, I find, do group their books by subject matter (philosophy, gardening, health, sports, and so on).


I’m a bit old-fashioned (alas, I have not read Harry Potter, and I have already built up my collection of music on physical media), and I have to confess that even with my computer music files, I cannot help but organize them into subdirectories by composer, instrument, and form (concerto, symphony, etc.) at the very least. I still find the “album” concept convenient to record concerts, for instance – I find that the 1-hour format of most media (LPs mostly 40 to 50 minutes taking both the sides) is tailored to the average length of most performances. So even if I do manage to digitize all of them (or procure digital versions), I think I will still organize the files on my hard disk in a proper subdirectory structure, and I will probably follow a formal classification scheme like the Dewey Decimal to do so. And I will preserve the album cases and covers (especially the old LPs) for their erudite notes, beautiful graphics and illustrations, and the way they are evocative of places and events that computer files just cannot match! I’ll have occasion to describe my experiences with some of these computer classifying and cataloguing packages in future posts.

17 Classifying recorded music with Dewey

Actual music comes in various media – tapes, plates (LPs, for example), discs of various types (laserdiscs, audio CDs, DVDs), and so on. I don’t think anybody would think of mixing these objects with books on the shelves – they will collect dust, and be of different sizes and shapes from books, that will call for different handling. So the actual physical media tend to get stored in separate locations, probably under a shutter  or door of glass or other material.

My usual approach to such classification issues is to first visualize where I would put them normally. In this case, I’m pretty sure that I would like to stack the LPs separately, singles separately, the tapes separately (by size, but I have only micro-cassettes), then CDs (and DVDs and video discs with them, probably). Within each type, I’d probably arrange them in the standard Dewey Decimal order for 580 Music, just as if they were books (treatises, texts). All that remains is to give some mark or tag to show what type of recording media each item is. A simple way, obviously, is to prefix each number with a code symbolizing the type: LP (Long Play platter), EP (Extended Play), SP (Short Play), ACD (Audio CD), VCD (Video ditto), MC (Micro Cassette) or CC (Compact ditto), DVD, MP3, and whatever else you want. Of course, this will split a particular performer’s works among a number of locations or catalogues, so if we wish to keep them together, we could add the type of physical media (LP etc.) after the Dewey number and performer, so that a mechanical listing (by a computer, for instance) would list a particular performer’s ACDs, then LPs, and so on. 

For Western music, it is usually convenient to classify by genre and instrument (represented by the appropriate Dewey number) and composer, represented by the initial letters of the name, then the musical form (if desired), year, and serial number, if needed. Of course things can’t be always simple, and some “local” innovation may be called for to group symphonies together, or violin concertos together, and so on. For Hindustani classical, it’s usually the instrument that is the distinguishing facet, then the performer (not the composer), then year. Since each item may have pieces in a number of genres, it may not be so important to specify this in the classification number; the manufacturer’s name and the item’s serial number may be more useful to distinguish similar pieces. Enough letters would have to be carried for names to distinguish them clearly. For Western names, the surname is usually the entry point (Beethoven, rather than Ludwig, although the Bachs would need both the family name and the individual’s names). For Indian names, the surname is usually boring, because, like the old king who gave each of his three daughters half his kingdom, half the names are Kumar, half are Singh, and the remaining half are Khan (more or less!); I find it much better to enter with the first name, e,g, Ali Akbar, Allauddin, Rashid, and so on for the Khans. Thus an audio CD of a vocal recital by the Hindustani classical singer Rashid Khan would be ACD-789.9H’1’32 (for solo voice) followed by RAS 2011, and maybe the serial number. Or if I wanted a combined list for all media by the artist, 789.9H’1’32 RAS 2011 ACD, 789.9H’1’32 RAS 2010 DVD, and so on (this is a purely local innovation, not standard as per DDC!).  

Dewey declares under 580 that it “does not distinguish scores, texts, or recordings”, but goes on right thereafter (in the usual delightfully contradictory style we have come to love) to offer a choice of three methods of doing so.  One is to prefix a letter or other symbol, such as R for “Recording”, M for scores, etc. to the usual Dewey number for a treatise (which is the first method illustrated above). The understanding is that one goes to the appropriate storing place for each type, say the “Recordings Room” for R’s. In my institute’s library, they have put all the annual reports, project documents, and such like, in a separate room, and the catalogues show this location by the prefix D for Docs. Thus, Beethoven’s violin concerto could be classified as R- 787.2 (for Violin), followed by ’1’86 (for Concerto from 784.186), followed by composer, giving say R787.2’1’86 BEE 1964.  Of course, this would scatter Beethoven’s works all over the shelves, so to keep each person’s works in one place, we may have to alter the order in which these elements are entered: R-BEE-787.2’1’86 OIS (for Oistrakh, the violinist) 1964 (a rather non-standard way of achieving it!).
The second method provided by Dewey to segregate recordings is to add to the number for texts, the numbers following 78 in the range 780.26-780.269. As mentioned earlier, standard subdivisions of 780 Music are modified in places to cater to the special requirements of the subject. 780.26 is actually 78 with the standard subdivision -026, which in Table 1 is actually Law (but not recommended for developing numbers, preferring to use the main numbers 341-347). Under 780 Music, however, the standard subdivision -026 (actually, -26, as the zero is already provided by the base number 780) and its further subdivisions are used for a different purpose: “Texts, treatises on music scores and recordings”. Under this, then, 780.266 is “Sound recordings of music”. The number, when used normally, would refer to treatises about recordings (like the various guides to recorded music),  but Dewey is suggesting that we use the latter part of these numbers for the recordings themselves, or for the scores: 787.2’0266, recordings of violin music. This standard subdivision -026 can be used wherever an “add as instructed” from 780.1-780.9 is provided: thus, 787.2’1’86 (for Violin Concerto) followed by ’0266 (for Recordings), 787.2’1’86’0’266 BEE 1964 and so on, neat! Standard subdivision -0267 likewise referes to “Video recordings of music”.  (As far as can be made out, we have the option of adding suffixes through connectors -1- or -0- any number of times).

The third option suggested by Dewey is to class recordings under 789, and instructions at that number suggest using an alphabetic mark for composer, followed by the numbers after 78 in the range 780-788.


I must confess that I have not actually gotten round to classifying my recorded music under Dewey or other system. What I have is a list of these items (LPs, cassettes, CDs etc.) grouped by composer in the case of Western classical, and by performer in the case of Hindustani music. This is maintained physically in a loose-leaf ring binder of half the normal page size, so that pages can be added for new names or items as needed. The same information is also entered in a computer database (I use Lotus Agenda® about which I will talk later), which is based on DOS, and has never been ported to the Windows environment, alas! Since many of these albums (as they are technically called) combine say concertos and sonatas, or Hindustani khayal and thumri, and so on, there is not much scope for following strictly the Dewey order of musical forms; however, I broadly class vocal forms first, followed by the main instruments in the Dewey order. Mixed albums, of course, are located in front (or top). The lot are kept in various shoe boxes (ideal for CDs!) arranged alphabetically (by first name of artist for Hindustani, standard family name of composer for Western), and the LPs, of course, are stacked in a cupboard.