Showing posts with label Classifying. Show all posts
Showing posts with label Classifying. Show all posts

Tuesday, December 2, 2014

19 Classify and Catalogue – two sides of a page

One small point that may be worth pointing out here is the distinction between two parts of the process: one is cataloguing, the other is classifying.

A catalogue (catalog) is basically a list of all items: they could be our own possessions, or they could be even a seller’s stock list (or inventory) or our own wish list. When it comes to books, there are usually two basic types of lists used: one is an Author Catalogue, which starts each item with the author’s name, followed by whatever details we feel are needed to identify the item uniquely. The other type of list is a Subject catalogue, where the first entry is the code number or word for the subject. If anybody remembers, our institutional libraries used to have these two types of catalogues written on cards of approximately 3 by 5 inches, stored in sliding trays that formed an impressive piece of furniture near the entrance. Earnest scholars would spend hours thumbing through these cards, trying to locate their particular requirements.

Both types of catalogues have their uses. If we wish to locate records (cards) for books by a particular author, say Dickens, and we do not know where in the library shelves these books will be found, we go the Author Catalogue and pull out the tray for the D’s. Of course, the utility of these cards depends on what else is recorded on each book’s record: usually it includes the serial number of the book in the library’s stock register (the Accession number) which should be a unique identifier, the book title, year of publication, and publisher’s name and edition, and perhaps the international standard book number (ISBN). There are detailed codes written for this (the Anglo-American Cataloguing Rules AACR2, for instance) But apart from these, the most useful to locate the book would be a location number or code. In the least ordered library, they could be simply stacked in the order of their accession numbers, as we suggested could be done for reprints of papers, as they don’t have solid spines which will enable them to stand up on their own in the shelves. For most collections, however, it will be nice to have them grouped by subject, which is where the Classification scheme comes in.

Commercial bookshops usually do not adhere to a very strict code of classification, and generally group books by broad subjects, like Physics, Chemistry, Sociology, Politics, History, Current Affairs, and so on, maybe even under further subdivisions if it is a campus bookshop, say different branches of Chemistry or Physics or whatever (usually following the university syllabus). Meticulous (let’s face it: somewhat obsessive-compulsive) documentation experts like Dewey (or to a greater degree, Ranganathan of the Colon Classification) like to reduce these non-standard subject headings to a standard code with a consistent system of labeling. Dewey, of course, uses mainly numbers: the 3-digit numbers stand for the one thousand subject headings (Sections), which are then expanded by adding on further digits to the right after a decimal point (the successive digits show the hierarchical position, rather than a value). The Colon Classification system follows a different philosophy, which I won’t even try presenting here (maybe another day!). It is my impression that the Dewey Decimal system is more popular because it has a strong management backup with the OCLC Online Computer Library Centre, Inc. (obviously), is constantly being developed by specialists at the Library of Congress, and more than anything has an infinitely more friendly and common-sense approach as against the Colon’s (let us face it) somewhat dry language, exaggeratedly punctilious rules and cryptic terminology (well, one has to admit that it lives up to its rather unfortunate name).


The Dewey number locates the book under the broad discipline (more or less the 3-digit Section heads, but also maybe under further subdivisions where there are distinct areas), and then under the most appropriate specific subject according to the entries in the schedules. However, for large collections, one can go even further, by adding various suffixes from the half-dozen Tables provided in Volume I. These enable various aspects or facets to be specified: a favourite is the geographical or regional coverage, for instance, from Table 2. The Standard Subdivisions in Table 1 provide tags for various aspects: the suffix -01, “Philosophy and theory”, for example, is the first subdivision provided in many numbers in the schedules (it could be used for Policy as well). The suffix -09 introduces the geographical locations, persons, and periods from Table 2, and so on. Table 3A and 3B provide tags for different numbers of authors and so on (for collections and anthologies, for instance). Table 3C provides “Additional Notation for Arts and Literature” to be added “where instructed”. Table 4 provides “Subdivisions of Individual Languages and Language Families” (from 400), Table 5 “Ethnic and National Groups”, Table 6  provides tags for “Languages”, while Table 7 for types of persons has been deleted.

Friday, November 28, 2014

18 Why classify when “Everything is miscellaneous”?

While on this topic of classifying our music resources, it would be as well to revisit the question of why at all we want to have a classified list of our possessions. Why not just keep a running list where we enter each thing as it comes, something like a “general ledger account” of day-to-day transactions?

I came upon a very interesting book on this question (yes, there are geeks who write whole books on as mundane an activity as classifying and arranging!) titled “Everything is Miscellaneous – The Power of the New Digital Disorder”, by David Weinberger (published 2007 in Times Books by Henry Holt & Company, New York, ISBN 978-0-8050-8043-0). Weinberger is described as a fellow of the Harvard Law School’s Berkman Center for the Internet & Society, and an adviser and consultant for Fortune 500 companies, bestselling author, and a doctor of philosophy, so he should know a thing or two about the subject. His thesis is that the power of the computer and the Internet have placed huge databases at the call of a button, and searches on key words can be made in a fraction of  a second, so nothing really needs to be classified any more. In other words, everything can be entered in a single, massive list of all things.

The prime example Weinberger cites is, appositely, from the world of online music resources, the Apple iTunes music store, where the albums I referred to in the previous post, have been unpacked as tracks, enabling consumers to download just what they want, when they want, thereby giving Apple “more than 70 percent of the market”. The files need not be organized on your hard disk in any sort of structure, as the software allows you to search on key words any time (it will be faster, I suppose, if all the filenames and their embedded metadata (data about data!) are already indexed on every word, like Google does for the entire Internet). By giving keyword filters, we can narrow down the choice as much as we wish. Even filenames need not be intelligible (say Western-Bach-Concerto-Violin-No.1), but can just be a miscellaneous number, as the file’s embedded metadata would have information on which searches and selections can be made.

Another example Weinberger likes is the way digital cameras have generated millions and millions of pictures that are stored on the Web, viewed and exchanged and commented upon on the Web, and seldom printed out or put into albums. Here again, programs like Flickr.com have relieved the consumer or average user of the responsibility of classifying the pictures – as long as some metadata is embedded to indicate date, or maybe location, even the subject may not be all that relevant. What Weinberger says is that in order to “take full advantage of the digital opportunity – we have to get rid of the idea that there is a best way of organizing the world”.

On the other hand, if you tried this with physical objects – your collection of photographic prints or slides, or your books, or your CDs – you would end up with a big pile of stuff that you would have no way of using purposefully unless you did some sort of sorting – usually by subject, not colour or size! There actually was such a bookshop in Bangalore, where I live, which was more or less an icon – the owner could locate almost any book in his pile, but few others could. The store, sadly, closed a few years back, but another phenomenon has sprung up in many cities across the country – shops selling huge collections of used books imported by the carton (we are talking of shipping, not cardboard), sometimes even sold by weight! Most of these stores, I find, do group their books by subject matter (philosophy, gardening, health, sports, and so on).


I’m a bit old-fashioned (alas, I have not read Harry Potter, and I have already built up my collection of music on physical media), and I have to confess that even with my computer music files, I cannot help but organize them into subdirectories by composer, instrument, and form (concerto, symphony, etc.) at the very least. I still find the “album” concept convenient to record concerts, for instance – I find that the 1-hour format of most media (LPs mostly 40 to 50 minutes taking both the sides) is tailored to the average length of most performances. So even if I do manage to digitize all of them (or procure digital versions), I think I will still organize the files on my hard disk in a proper subdirectory structure, and I will probably follow a formal classification scheme like the Dewey Decimal to do so. And I will preserve the album cases and covers (especially the old LPs) for their erudite notes, beautiful graphics and illustrations, and the way they are evocative of places and events that computer files just cannot match! I’ll have occasion to describe my experiences with some of these computer classifying and cataloguing packages in future posts.

Sunday, March 18, 2012

05 The DDC thousand Sections (subject heads)

The DDC thousand Sections are on their own page (see the tabs on the top of the screen).
The following links give the Dewey summaries:

The summaries are also available on Wikipedia:
http://en.wikipedia.org/wiki/List_of_Dewey_Decimal_classes

Life can never be as simple as it looks, and as far as the DDC is concerned, there are revisions every couple of years which needs a fresh ‘release’ or version. I myself still use the DDC 20 version, which mainly had a major revision of the Music sections. There have been quite a few major changes since then, especially in Computer subjects (000s) and in the Natural Sciences (500s). The latest is the 23rd print version, DDC23 released in 2011; the following list (taken from the DDC website) refers to the DDC 22 version released in mid-2003; DDC 21 was released in 1996 (more of versions in a later post). All copyright rights in the Dewey Decimal Classification system are owned by OCLC. Dewey, Dewey Decimal Classification, DDC, OCLC and WebDewey are registered trademarks of OCLC, oclc@oclc.org.

The following links will take you to downloadable pdf documents from the OCLC site, providing an introduction to the DDC 22, a glossary, and a guide to the major changes from DDC 21:
The following links will take you to similar pdf’s for DDC 23:
Here’s a link to the OCLC blog: http://ddc.typepad.com/025431/

In my experience, classification can become a puzzling affair if taken to the extreme; one has to draw the line somewhere comfortable, and choose whatever number gives a close enough fit until further research is feasible (or look it up in their blog!). There are ways to cheat a little, like looking up your book (if it’s a published one, or else a close alternative with a similar title) on any of the public library websites; my favourite is the British Library (http://www.bl.uk/)!
Apparently, the numbers up to the Sections (3 digits, before the decimal point) are freely available on public media, but the detailed classification scheme to the right of the decimal point, would require you to purchase a copy of the manual (or its web version, of course).

04 The DDC hundred Divisions

Before we go on to list the thousand Dewey subject numbers and subject headings, let’s take a look at the broader structure. The DDC groups all subjects in TEN broad CLASSes, each starting with a round ‘hundreds’ number:

000      Generalities (Computer Science, Information, and General Works)
100      Philosophy and Psychology
200      Religion
300      Social Sciences
400      Language (and Linguistics)
500      Natural Sciences and Mathematics
600      Technology (Applied Sciences)
700      Arts (and Recreation)
800      Literature800 Literature
900      Geography and History (and Biography)

So these are the ‘Hundreds’. All the 900s – that means the numbers 900 to 999 - cover the field of Geography and History, and so on. Each of the Hundreds Classes is subdivided into ten DIVISIONS each (900 to 909, 910 to 919, and so on), each Division into ten SECTIONS. Ten Classes, into ten Divisions, into ten Sections each: totalling up to 10 X 10 X 10 = a 1000 Section numbers, a thousand subject heads.

I've put the hundred Divisions on their own Page...see the tabs at the top of the screen.
So where does the ‘Decimal’ in the DDC come in? That’s because the subdividing doesn’t stop here; the process goes right on, except that a decimal point is put after the 3-digit number. Thus 910 can be subdivided into ten sub-classes (910.0 to 910.9), each of these could be further divided into ten sub-sub-classes (910.90 to 910.99), and so on… to as many levels as we wanted. Strictly, the numbers ending in zero after the decimal point (like 910.0) are not given separate mention, so the total number of subdivisions may be 9, not 10 for each level. Each number is associated with an individual subclass of the main subject. This is usually set down in detail for each number, as each subject will be broken down in a domain-specific manner, but there are some broad conventions, such as .01 refers to the ‘philosophy’ or ‘theory’ of the main subject. There is in fact a Table of Standard Subdivisions, which we will describe later.

03 The Dewey Decimal scheme: playing by numbers

Now we can address the core problem, which is: how to arrange our non-fiction (the ‘subject-matter’) books in a meaningful manner. The basic premise is, of course, that we can identify one main subject of each book or report in our hands. Dewey has done half our job by listing out a thousand main subject names, starting not from 1, but 000, to 999, in a remarkably prescient manner which conforms neatly to the usual practice in our own computer age (all the numbers are 3-digit ones, hence can perfectly fitted into a fixed field in a database structure).


For the first round of classification, perhaps this is all that we will need to arrange our books subject-wise. Within each subject, naturally we will be arranging our books by the alphabetical ordering of author last-names, thus (Richard) Dawkins would come after (Charles) Darwin. If you had two books by Dawkins, you would arrange them by their date of (first) publication. These three fields: DDC three-digit code number, Author name, and Year, would suffice to order your books in a perfectly predictable sequence.

The whole structure depends on the expectation that any new subject could be ‘adjusted’ under these thousand heads. If you did have a subject that didn’t fit into a single one of them, of course you’d have a problem…then there would have to be a sufficient number of ‘empty’ numbers to cater to these eventualities. In practice, most subjects could be fitted as a sub-category under the thousand main categories provided by Dewey.

One great feature of this system is that if you walked into any library or looked up any catalogue that also followed the same classification scheme (the DDC in this case), you would expect to find your favourite books and authors in the same sequence. If you were interested in say Physics, you would zero in on the 500s; if in History and Geography, on the 900s. These are the Dewey 'hundreds', which bring together related subjects in groups. A little down the line, you will be glibly talking in numbers rather than in names. Unlike bookshops, where each manager devices an individual scheme of arranging the subjects, in academic and public libraries, all of them follow the same. This is spoilt to some extent by the fact that there are two or three main systems; fortunately, the second most common, the Universal Decimal Classification (UDC) closely follows the DDC, and the hundreds numbers are all quite similar.

02 Let’s first get some simple schemes out of the way…

Let’s first get over some elementary, basic approaches to classifying and cataloguing your books. I have a certain brilliant friend, a genuine manager bureaucrat, who once was in charge of a premiere research institution, who wanted to know why they didn’t just ‘colour-code’ the books in the library and be done with it. He may have been kinaesthetic, or panaesthetic, or any of a number of novel, unusual ways of dealing with sensory stimuli. But for large collections, we will usually have to go for fairly structured approaches that depend on the most convenient pattern of arrangement for us in our workaday lives.

For a small personal collection, often nothing more elaborate may be called for than a simple Fiction/Non-fiction divide.  The Fiction is the easier portion, as most people don’t really want to divide it up into sub-categories, unless it’s by nation and language (English, English-American, French, Russian, German and so forth). Since the Indo-Soviet culture centres used to distribute amazingly economical volumes (cheap is not a nice word to use for them), I happen to have a middling collection of old Russians like Pushkin and Chekhov. If you have a large collection of, say, English literature, maybe you would like to group them by what is known as ‘genre’, like prose, drama, poetry, criticism, essays, and so on. Or, you could group them by periods… Ancient, Classical, Romantic, Nationalistic, modern, post-modern, and so on (I don’t know much about this, perhaps you would have specific classes in each nationality’s literature, depending on the watershed events in their history and evolution). Within such a category or sub-category, you would group them by author’s name, like Shakespeare, Wordsworth and so on (usually in the order of the surnames or family names, not by the first names, although we will come back to this in a later post). You would have one set for classical works (Literature or belles-lettres as it’s termed), and another separate series for modern fiction (popular stuff, pulp, romances, and so on). The dividing line is a bit vague…where would you put Agatha Christie, for instance. Maybe you could just do post-World War II and pre-WWII and be done with it.

This is more or less the scheme in the Dewey system, too. The same could be extended to all Non-fiction as a whole, if you have very few in this category. Personally, I feel that sooner or later this will become too limiting, so I would much prefer to start dividing them up by at least some major subject categories right from the beginning…at the minimum, say Humanities and Sciences. Within each, of course, you would arrange the books in alphabetical order of the author names.

There are some other plausible criteria for dividing your collection. For instance, big hard-bound picture books and encyclopaedias could go into a shelf of their own, within which they would of course be arranged in the order of the author names. In libraries, this may be called the ‘Folio’ section referring to the big size, which anyway calls for special shelving, or maybe the ‘Reference’ section to denote their high value, so this is not as daft as it may sound.

Another scheme would be to separate His, Hers, and the Kids’ books. That will avoid recriminations. In fact, I would strongly urge each person to only fiddle around with their own collections, and not touch their spouses’ or their kids’ books and records…you have been warned.

Saturday, March 17, 2012

01 What this is about

How do we manage our acquisitions of books, papers, magazines, tapes, CDs and other things we just can't bring ourselves to give away? These are the things that in a way define our life experience, our travel through the world. They've kept us company through happy days and sad, through exciting times and dull periods. But left to themselves, they become almost useless as they are all jumbled up, they get tucked away in dusty corners and tattered cardboard boxes in the basement, and they introduce health-threatening clutter and disorder.

This question may not loom very large at the beginning of our acqisitorial lives, but a few years down the line, a few shifts of residence, and the issue of keeping, storing, and retrieving them at our will and convenience, becomes important.

My own response was to take the bull by the horns, grasp the nettle, and organize my books and music using recognized classification schemes like the Dewey Decimal Classification (DDC), although there are other... especially the one I started with, the Universal Decimal scheme (UDC), and for Forestry (my profession), a modification of it called the Oxford (ODC). Let me share my experiences in the hope that it may be of some help to some blighted soul groaning under the weight of all that rubbish.... but organized even imperfectly, can be enjoyed and used by self and others...