Showing posts with label digitisation. Show all posts
Showing posts with label digitisation. Show all posts

Monday, 28 March 2016

Digitisation: practicalities, pleasures and pitfalls. #HistLib15

On 12 November last year I went to the Historic Libraries Forum annual conference, titled ‘Practicalities, pleasures and pitfalls: uncovering digitisation projects’. As ever, it was a really informative and useful day, and a great chance to catch up with people.

The theme of the day was digitisation: how to plan, run and maintain a successful project. As someone who'd been involved for the last 18 months as one of the partners in the UK Medical Heritage Library project, I could only wish that the conference had come round two years earlier: it would have been a great help! But no matter, the case studies and advice generously shared by the six speakers will certainly be of use as I plan projects in the future.


1. Calum Dow, Towsweb Archiving. ‘The importance of planning for digitisation’

Townsweb is a company that undertakes digitisation for individuals and institutions. Calum provided a very clear overview of what you need to consider when considering and planning a project.

Key points
  • Start with formal goals: use SMART objectives. Consider your desired outcomes, and the purposes of the digitisation: is it for for preservation, for display, for searchability, for access, for revenue? Who are the core users of the digital collection? What are your ideal timescales and milestones?
  • Communicate with your supplier! Make sure everyone understands your project goals, including the end purpose of the digitisation. Ensure your supplier is happy to keep you updated.
  • What to digitise? You might choose the items in greatest use, or at greatest risk. Ideal concerns have to be balanced with the items’ suitability in terms of condition and format.
  • Image outputs and resolutions: capture your master images in an open-source, non-proprietary format.
  • Access: metadata is vital to both your internal archive or document management system and your external website.
  • Practicalities: you’ll need to consider how and where digitisation will take place. Different formats and materials have different requirements. The better quality digitisation will usually be achieved offsite. Very rare or fragile items may have to be digitised on-site, but you need a room with controllable lighting, good power supplies, space, etc.
Resources

2. Damian Nicolaou, Wellcome Library. ‘Hand in hand: The ups and downs of partnerships in digitisation’

Damian presented several case studies of collaborative digitisation projects. There can often be encouragement or inducement to work collaboratively these days, either with commercial or non-commercial partners. The Wellcome Library works with lots of partners. There are commercial (eg ProQuest) and non-commercial (eg the Internet Archive) organisations doing the digitisation. And then there are the organisations that use the output (eg Zooniverse), funders (eg JISC), publishers, academic partners, software developers, licensing societies, and internal partners. Some of the projects are London’s Pulse, ProQuest Early European Books, and the UK Medical Heritage Library.

Various themes recurred throughout the case studies that Damian presented:
  • The difficulties of balancing your priorities against your partner’s (or partners’) priorities.
  • Are the people at the coal-face of the digitisation working directly for you or for someone else?
  • How to communicate well in order to manage relationships and expectations. Be clear who has responsibility for what.
  • When collaborating, you need to agree tight standards for images and metadata. Sometimes the standards are baggy and imprecise, and even though everyone is adhering to them, local practice variations still make interoperability difficult.
  • That doesn’t only apply to metadata: agree everything to do with the project in writing and even then interpretation may differ
  • Things will change over time, including the people and staff you’re working with in partner organisations, and the goals of those organisations can change, too.
  • Can you live with the conditions imposed by (commercial) partners? Are you prepared to do things their way?
  • Could you do it on your own?
  • Even ‘free’ things have costs.
Aside from the points directly relevant to collaborative working, there were other points to consider for all projects.
  • Consider use cases for the digitised material both for the present and future, before you sign: keep things open!
  • Broaden your reach by hosting the digitised content in multiple places (eg Wellcome Library website and the Internet Archive)
  • Don’t neglect your internal departments, eg the conservation team.

3. Abby Matthews, Sutton Archives. ‘From Cellar to Hard-Drive: The Challenges of Digitising Glass Plate Negatives’

On a much smaller scale than the Wellcome Library, Sutton Archives is digitising fragile glass-plate negatives. ‘The Past on Glass’ is a 2-year Heritage Lottery Fund project, working with negatives from a photography business that was abandoned in 1918. The collection of negatives is vast, and there were several issues to address.
  • The time-frame of project was set in advance, at the point of applying for funding, with little idea whether it was realistic or not, so part of the project has been working out how to best prioritise the work.
  • There is a large and unknown number of damaged plates, so they’ve prioritised those in good condition.
  • The project relies on volunteers, who have varying skill sets, so the workflows need to be written very carefully, with different volunteer roles available.
  • There are a limited amount of resources available, including space and equipment, so they have been inventive about how to make the most of the scanner.
Abby gave some really useful advice:
  • Do as much research as possible before you start. Seek advice. Know your options.
  • Think carefully about specific needs and long-term requirements.
  • Be prepared to go your own way.
  • Set strict project methodology and workflow, and make sure everyone follows it.

4. Jamie Robinson, John Rylands Library. ‘CHICC - Heritage Imaging at The John Rylands’

Jamie described the work of the The Centre for Heritage Imaging and Collection Care (CHICC), based at the University of Manchester Library. They work for the university and external organisations, too. With university items, they try to digitise the whole item when there’s an imaging request for one part of it, to save having to do further work again later on the same item. They are constantly updating their equipment, and make the most of new developments, for example working with a pigment specialist on how best to image gold used in printed books and manuscripts.

Key points
  • Different specialist materials need specialist imaging: this developing all the time.
  • New techniques, such as multi-spectral analysis, have a lot to offer, and are leading to new research avenues on the primary sources.
Key resources

5. Steven Archer, Parker Library, Corpus Christi College, Cambridge. ‘The legacy of Parker on the Web

Parker on the Web was one of the first major manuscripts digitisation projects. It launched in 2009. Steven spoke about how digitisation is only the beginning with a project of this scale and scope. The project has a large maintenance demand: corrections and additions to images, references and bibliographies appear in semi-annual releases. In addition to this, the interface is now starting to show its age to some extent, so plans are underway to launch a new version in due course.

Steven also had some interesting reflections on the way digitised content has changed reading room culture. Remote access gives the librarian much less idea what people are working on, and thus fewer outlets to help them! They hope that the new version of Parker on the Web might have more facilities for interaction between scholars working on the manuscripts.

Some resources mentioned:


6. Naomi Korn, Naomi Korn Copyright Consultancy. ‘Copyright issues in digitisation’

You can't consider digitisation without considering copyright. And copyright is complicated. Works whose authors died before 1969 and whose works weren’t published by 1 August 1989 are still in copyright until 2039. Something that looks like it might be 'one' item, like a scrapbook or diary, might have multiple copyright holders. Naomi had some very practical advice for how to manage rights in (our rights (or lack thereof) to digitise and reproduce) and rights out (what rights we’ll grant other people to use the digitised images).

Key points:
  1. Be aware what the digitisation contract says re: what’s expected and required. You need to exercise due diligence!
  2. Be aware of layers of rights: for example, cuttings stuck into a diary. Resources are needed to check and clear rights.
  3. Rights management is risk management.
  4. Check what the terms are regarding the licences on the digital images. Heritage Lottery Fund projects are required to use a CC-BY-NC licence, so you need to ask rights holders for this level of permission. When asking for rights you can ask for equivalent rights, or ask for as much as possible (but latter has risk that rights holders will say no as asking too much.)
Tools for reverse image searches (to try to uncover image sources)

Conclusion

So, to sum up… Knowledge is power when thinking about a digitisation project. The more detail and understanding you have about the materials you want to digitise, the easier the planning and implementation should be.
  • Think before you act! Know your collections. (Really know your collections.) Know your audience(s).
  • Think about the future: allow for possible new uses.
  • Host it far and wide – the more places, the better.
  • Have clear project plans and workflows.
  • Talk to everyone as much as possible, and don’t forget your internal partners and collaborators.
  • Plan for what happens after the project finishes: how do you maintain it?
  • Factor in the time and money involved in copyright clearance from the start.

Sunday, 16 November 2014

What are we here for and where are we heading? #dcdc14

At the end of October I went to ‘Discovering Collections, Discovering Communities’ (#dcdc14), a collaborative conference between The National Archives and Research Libraries UK, exploring ‘the ‘discoverability’ of collections across different formats, institutions and professions’. It had a focus on archive and museum collections and institutions, and was held at the new Library of Birmingham. The full programme is here (pdf link) and video of sessions will be going up in due course.

Birmingham Library by Chris(UK), on Flickr
Creative Commons Creative Commons Attribution-Noncommercial-No Derivative Works 2.0 Generic License   by  Chris(UK) 
There were 9 panel sessions and 10 workshops split over the two days. I went to:
  • Panel 1: Exploring the mechanics of cross-sector collaboration
  • Panel 6: Visualising the digital discovery
  • Panel 7: The community within: volunteers as a route to discovery
  • Panel 9: Collecting - for whose sake?
  • Jisc Workshop: Finding and Using Digital Collections: Do we need new tools?
There was a lot of interesting and engaged discussion on the Twitter hashtag, which cut across the different sessions and themes and helped to make connections between them. I think it’s also the first conference I’ve been to that had an end-of-day summary session in which each panel chair gave a short summary of their session. I found that very useful, too.

I found it interesting to hear about lots of different projects going on at institutions of different sizes, and in different sectors – archives, museum, and library. It’s easy to forget the different approaches of the three domains, given how often they’re lumped together as ‘heritage’ or ‘memory institutions’. It came out early on that museums have a very different attitude to research use compared to archives and libraries: it’s what the latter two exist to facilitate (in large part), but it’s not what a museum is set up for. So when a keen academic turns up to a tiny museum and spends a long time looking at artefacts or talking with the curator, that can feel like a huge investment of time for the museum, with little tangible reward.

Three broad themes appeared to me over the three days. So, as ever, I won’t précis each session I attended, I’ll throw together some thoughts under headings.

Representation

This is something that I think all of the first three panels touched upon, but which certainly struck me in a session that was ostensibly about cross-sector collaboration. One speaker mentioned that in a museum display of archival material exploring ‘hidden histories’, people to whom no photograph or artefact could be attached had to be omitted from the display because of the visual demands of displays.

In the social media panel there was discussion of the double-edged sword of high-volume but low-depth attention: can we call click-bait archives images that go viral (44 medieval beasts that cannot even handle it springs to mind) successful engagement? Or, under what circumstances can we call that successful engagement, and how to we balance the demands and delights of that kind of promotion with archival collecting and use that allows for and encourages the discovery of more complex, more nuanced, stories and understandings?

It’s a concern that digitisation, an increased focus on exhibitions, and/or a focus on social media overly privilege the visual and the tangible and disadvantage the textual? Lots of people talk about their exciting online collaborative projects exposing hidden histories and forgotten stories, but I worry there’s a risk that we do this at the expense of another, doubly disadvantaged set of histories that aren’t sufficiently photogenic.

I’ve pulled together a Storify of the Twitter discussion on this topic, rather than paraphrasing it all in this post. 

Strategy

Collaborating

The first session I attended was all about the ‘mechanics of cross-sector collaboration’. We heard about the Making Britain project – a collaboration involving the BL, the OU and many others – and the Inspiring Women project – a collaboration between Tunbridge Wells Museum and the University of Kent. They both sounded like really interesting projects, but in both cases the collaboration was born out of pre-existing relationships and networks: people already knew people who’d be interested in working on the project. Now, there’s nothing wrong with using your network, but it can seem like an impossible task to create this sort of collaboration if you don’t have a handy connection already.

The third paper in the session sought to address this. The Share Academy is a project run by UCL, the University of the Arts London and London Museums Group that aims to ‘build sustainable and mutually beneficial relationships between the higher education sector and specialist museums in London’. They strongly advised taking a more strategic approach to collaboration, by taking a pragmatic approach to planning. It’s important to have collaborative relationships documented and agreed at the start: what will the outputs of the project be (academics and museums might want very different things)? What happens if and when key people move to new institutions? How much investment of time, money and resources is expected from each side?

One really striking point for me was the comment that small museums can feel used by academics who come and research their objects and then publish without seeming to give anything back to the museum. There’s a really large investment of very limited staff time involved in providing access for researchers, and it seems that museums having necessarily been doing a great job of explaining that – unlike in libraries and archives – facilitating this access isn’t part of their core work, and that they would appreciate a more equitable relationship.

Unfortunately, Share Academy is only short-term funded, so it’s not clear whether it will continue, and be able to act as the matchmaker many of us would like!

Lots of people agreed that national advice and support on setting up and running collaborative projects would be welcome, and there’s promise of this coming from TNA. In the meatime there is, for example, this report on collaborative working practices in science heritage: ‘Mind the gap’ (link via Melinda Haunton).

Collecting

Karen Pierce presented a really inspiring paper about her work on the History of Human Genetics library at the University of Cardiff Library. This is a collection of materials concerning the history of the study of human genetics.

It was conceived of by a geneticist – Peter Harper – who was concerned that the history of this comparatively new discipline wasn’t being preserved. It now includes three complete personal libraries as well as selected donations from other people’s collections. It includes printed books – including classic textbooks and other works – as well as grey literature. There was a conscious decision to collect grey literature, because it represents an otherwise undocumented stage in research: networks, transient developments, events, suppliers and so on.

(The quote of the conference was undoubtedly Karen’s two descriptions of grey literature: 'publications by organisations whose main business isn't publishing' or 'floppy stuff’. The subsequent cries of recognition and anguish from cataloguers on Twitter is worth a read.)

This paper really made me think how fine the line is, and how variable is the location of the line, between a library’s perception of something as ‘usefully collected stuff from and expert’ and ‘endless shelves of rubbish we don’t need’. There’s real, serious, professional skill in evaluating the value (current and potential) of collections particularly of this type, and sometimes the smart professional decision is to say ‘we don’t need it’, but this isn’t always so. Sometimes we need to be brave enough to say ‘yes, we’ll take it, and love it, and make it great’.

Karen’s paper was an inspiring example of how to negotiate this boundary by setting clear aims, remits and procedures for such a collection. It’s great to see that at least one place is negotiating this successfully, rather than succumbing to the inertia of the ‘bay and a half of Stuff that someone gave us three librarians ago’.

Data and the catalogue(s)

Ah, shouldn’t we have got past worrying about catalogues by now? More than once people said and tweeted ‘the aggregator is dead’ and ‘it’s all discovery now’ and so on, but in the JISC workshop on digital discovery tools I stuck my head above the parapet and pointed out that the one key thing that would improve the discoverability of my (so-far-non-digital-)collections would be better catalogue data. Or better metadata if you prefer. It turns out that I wasn’t, in fact, expressing the view of a behind-the-times old fuddy-duddy, but actually a feeling that was in the room and online, too. Never mind aggregators being dead, we’re still trying to sort out how to get stuff into them nicely.



Elsewhere, it was pointed out that even if we have all the resources to record all the stuff we want to record, we don’t have cataloguing systems that really allow us to usefully (or at all) record the stuff that makes special collections well, err, special.

A catalogue record for a Shakespeare first folio? Quite possibly won’t mention that it is the ‘first folio’. Probably can’t easily link through to that amazing exhibition you had about it last year. May manage to link through to the digitised version you have, though not necessarily. Unlikely to record other context or importance... As curatorial tools, and tools for the people Out There to understand our stuff, there’s a lot that’s still desired.

(Incidentally, the problems of what’s in the catalogue not matching what the users are looking for also cropped up a recent event at Harvard. We can’t easily meet the needs of researcher’s who’re interested in all sorts of copy-specific features, because they just haven’t been recorded. This extensive conversation is well worth a read – it brings out several different issues affecting our ability to achieve this.)

In summary...

There’s so much going on out there, and there’s so much that we can all do in our services (big or small). There’s huge pressure to do more, to reach more people, to work with more people, to get out stuff out there as much as we can. This is facilitated by things like social media, the spread of digitisation projects, and so on, but these tools don’t themselves ensure that we’re doing things well. We need to keep in sight that we should be doing stuff for a reason and that we ought to be able to set and measure against criteria for for doing this stuff well would look like and we need to consider how it will be sustained in the future.

Sunday, 10 November 2013

Making the most of possibilities of digitisation: an #rbscg13 write-up

Rare Books and Special Collections Group Conference (programme (.docx)).  The theme this year was digitisation (last year was fundraising and advocacy).  The three recurring themes across all of the papers were audience, metadata, and the 'thinginess of things', with some very useful practical advice thrown in. I'm not going to give a blow-by-blow account of every paper, but rather to pick out the bits I've been chewing over since.

All three themes were addressed in Simon Tanner's keynote, which adeptly summed up the current situation as well as challenging us to think more imaginatively about the future and shaking us out of complacency about what we're doing now.

Tanner started by using the analogy of an ant mill to describe the fate of too many digitisation projects.  He emphasised that it's vital to plan properly and to seriously consider audience, opportunity costs (what aren't you doing if you are doing digitisation instead?), and to keep asking the 'so what' question to keep you focussed on why you're doing what you're doing. He also shared lots of useful resources:
Sian Prosser presented a case study on cataloguing and digitising a comparatively small collection of  manuscript fragments. She emphasised that even though there weren't many fragments, a very high level of specialist knowledge was required to describe them well, and that any project of this type is likely to take more time than envisaged.
  • TEI by example. Sian's project used TEI to mark up the descriptions. TEI by example is a set of free online tutorials.
  • Ransom Center Fragments. This is a very useful Flickr site displaying images from a large collection of manuscript fragments at the Harry Ransom Center, University of Texas at Austin.
My favourite paper from the conference was  Rowena Willard-Wright on 'Transforming our data for the internet on a tight budget'.  Willard-Wright works for English Heritage and described their work on digitising and improving their catalogue, including photographing objects, improving old records, cataloguing from scratch, and create varied means of access to the records.  Her talk really exposed infuriating and frustrating the problems faced by all cataloguers are.  As she spoke I wrote:
Willard-Wright's description of migrating and updating catalogue data is very familiar: data has been lost and garbled in transfers over years. Cataloguing has been and still is viewed as archane, and a thing not worth funding, because catalogues were and are seen as not for general consumption. I.e. they are perceived as been the exact opposite of their whole point. This problem that has been seen with the English Heritage catalogue is *exactly* what is being exposed as libraries moves from traditional OPACs to resource discovery/next-generation systems. And, most infuriatingly, the things that cataloguers have known and have been saying forever (e.g. consistency matters, access points matter) is suddenly being "discovered" as if it's new.
Willard-Wright was talking from the perspective of a museum catalogue, which is in some ways very different to a library catalogue.  Museums don't have such a tradition of the publicly accessible comprehensive catalogue, and write much more descriptive and less codified entries for their objects.

The English Heritage cataloguing project used teams of volunteers with very well-defined tasks.  They write clear, concise, engaging, small chunks of description - i.e. entries that confirm to the principles of good writing for the web. The volunteers aren't necessarily experts on the objects, and they're writing for audiences who aren't necessarily experts either. However, there's a recognition that the audience may have additional knowledge or stories to share, and for this reason a 'tell the curators something about this' button is being built into the public catalogue.  I absolutely love this - it's baffled me for years that so few library catalogues have a 'tell us if there's a mistake' button. Copac is a notable exception.  I fear that many libraries don't have one because, if it was ever mentioned in a meeting, someone piped up and said "but think of all the extra work" which likely trumped "think of how handy that will be for our readers, and how useful for us to make use of their knowledge".

That lack of connection to the audience was hammered home for me in another way throughout Willard-Wright's talk.  The museum descriptions are being written for general audiences.  Rare books records contain descriptions that are, frankly, written for librarians.  Not even, really, for most researchers. Yes, we include all sorts of useful information, but we code it up in impenetrable ways, and there's all sorts of information we don't include accessibly.  This has maybe been less of an issue in the past, when catalogue records were only seen by those initiated into our arcane world. But now catalogue records go along with beautiful/intriguing/important digitised books that all sorts of people might want to see, and our gibberish means *nothing*, and doesn't explain any of the basics. (How many records for the first folio show clearly that this is a first folio? Or the Nuremberg Chronicle?)

During Willard-Wright's paper Jill Dye commented that "The only difference between an online catalogue and a digitisation project is adding a photo?", and I think that in one way she's right: it's completely wrong to think that a digitisation project stands apart from cataloguing. However, making materials accessible in any way, but especially if they're freely available online demands a new attitude to description. We really need to step up our game.

Melissa Terras is director of the UCL Centre for Digital Humanities, and she presented a wide-ranging paper highlighting some of the possibilities of high-end digital imaging. She works on these projects in collaboration with computer scientists and engineers, and they're often adapting techniques already used elsewhere (such as in medical imaging). As well as drawing our attention to current projects, Terras made some important points about the theory and practice. Digitisation means lots more sorts of metadata needed so that we can properly interpret the images. As Hannah Thomas put it, Terras' work is "not just about creating a surrogate but about using the image to discover new things, inspire new research". Terras made the point that digital images are *not* exact reproductions of originals. Terras asked us to talk to her if we know of collections or items that would benefit from advanced digitisation and imaging work; part of her role is to connect the various people involved.
  • The one resource I'll share is this breathtaking (there were audible gasps in the room) video of the digital flattening of the great parchment book. Watch it. It's amazing.
Alixe Bovey spoke from the academic's perspective, and addressed some of the threats posed by digitisation. She was heavily involved in the campaign to try to prevent the sale of some of the Mendham Collection books last year. Bovey passionately explained that the digital is not the same as the physical, and we all need to communicate this better.  With the Mendham sale, the existence of digitised copies of the titles on databases such as EEBO and ECCO was used as justification for the sale, ignoring the copy specific details of the Mendham copies, as well as the failings of the scans themselves. Earlier digitisations have been particularly lacking. Never mind the poor black and white reproductions of scanned microfilm, they tended, for example, not to include any blank or apparently blank pages.in the source copy (see this post), and also ignored bindings, and made it difficult to determine the original size of the book. But we're not past such difficulties even with the best modern digitisations; they tend not to include scale rules, (see this post for difficulties of determining size), and give little indication of other factors such as weight, quality of materials used, or even smell.

Anne Welsh spoke very pragmatically from the point of view of libraries and library staff themselves.  She pointed out that we are continually needing to update and improve what we've done before: both content format and types of description.  She faced the fact that we can't do everything, and used the example of the University of Manchester Library Digitisation Strategy Group's 'Criteria for ensuring value to the Library for partnerships' (pdf link), which considers the value to the library of any potential projects.

Nicolas Pickwoad spoke about one element of early books which is too often overlooked in digitisations: book bindings. Most bookbinding digitisations (Pickwoad mentioned the Uppsala Probok project as an exception) show only beautiful, expensive, fancy and/or fine bindings, turning bookbinding digitisation has into "a decorative arts ghetto". This doesn't represent most early book bindings, which are less extravagant, but can tell us a very great deal about the book's history, and may often be the most interesting.

There's also a vicious circle at work: bindings aren't so often described in catalogue records, soscholars can't ask for them, so there's not so much research, so it's not seen as a priority... At least when books are viewed in person, the binding will be seen 'by accident' as it were.  If they're not included in digital surrogates they disappear altogether. Like many specialist aspects of digitisation, imaging bindings takes special requirements, including lighting, including to show structures accurately.
  • I'm keen to keep an eye on Pickwoad's Ligatus project, which is working on guidelines and terms for describing bindings better. It's hoping to develop vocabulary and multilevel descriptions for bookbinding including the ability to record negatives (e.g. 'no clasps'). This is key, because otherwise you just can't tell whether a feature is absent or it's just a bad record.

So all in all, my summary would be that we need to be using better, subtler and more flexible descriptive frameworks and presentational tools to make digitised materials accessible and available to the audiences who want to see them.  Digitisation can help with some, but not all, problems, and we need to advocate loudly for the intrinsic physical value of the things we want to digitise, to try to stem the tide of feeling that a copy is as good as, and entirely replaces, the original.
    There was lots of tweeting throughout the conference: