Thursday, July 3, 2008

Semantic Search Engine of African Languages

Here is a thoughtful article about the Kamusi Project from Appfrica: http://appfrica.net/blog/archives/82

The article talks in part about the blog widget we've been developing with one of our Code Africa volunteers, which people can insert in blogs and web pages to perform Kamusi lookups on their sites. The widget is not quite finished, but we put together the following brief presentation for a Barcamp event recently in Nairobi:


 


Thursday, May 8, 2008

Supporting languages that do not have localisation

Yesterday I had the privilege to present at a workshop in Milan for the ISO. The workshop discussed how ISO will continue its development in the 21th century. A whole day was filled with a mix of people inside and outside of ISO providing their point of view how the world is changing and many kinds of new technology are becoming available and relevant that have the potential to change the current practices at ISO.

Bob Sutor, the IBM vice president for Standards and Open Source opened and discussed everything from Wikis to Second Life. It was a great speech and it opened up the floor for the presenters that followed really well.

The WLDC is about languages and with Debbie's permission, she had seen my presentation ahead of time, I had included the WLDC as a way to establish that I am truly committed to do good for languages.. What we want to do in the WLDC is making a document languages and make a difference by doing this. To help us realise this, I approached Mr Sutor and asked him if IBM could be interested in giving languages a presence in the user interface provided by GNOME or KDE.

This is of a great practical importance; when you write Neapolitan for instance, you do not want an Italian spell checker telling you that what you have written is spelled is incorrectly. The localisation of software is an expensive and time consuming business, it is not realistic to expect that all languages, linguistic entities will be localised. It is however feasible to make Gnome or KDE aware of the language that is used for a document. This is the first step to ensure that this document will be tagged in its meta data appropriately to the language that is used.

I am sure that you know more great arguments why a practical application like this will be of a much bigger benefit then is immediately apparent. So please pitch in with suggestions so that we will be able to produce the proposal that Mr Sutor and IBM just cannot refuse :)
Thanks,
Gerard

Sunday, May 4, 2008

A proud moment

At the Wikimedia Foundation I have been banging the drum for the use of standards. I made some friends and enemies in that way, but the overall effect has been good. Some fights are no longer fought because the result is clear from the start.

At Betawiki, we are developing an extension for MediaWiki called Babel. The tool is to be used on the user pages indicating the self assessed skills in the languages a person knows. The texts are shown in the language itself.

When we do not have a translated text yet, we are still able to use the native name of that language courtesy of the data available in the CLDR. The standard is not complete, and I asked if it was possible to change the data in our database. I was told no. "The data belongs to a standard and, the data should be improved at source".

I do agree with this sentiment. I have written to someone active in the CLDR if there is an interest in collaboration. I am happy and proud of this turn of events. I hope that we are welcome :)
Thanks,
GerardM

Thursday, April 17, 2008

Of ancient and historical languages

According to the records at SIL the documentation for Ancient Greek (to 1453), ISO-639 code grc, has been tagged as type "Historical". This means that the language is dead. Latin lat on the other hand is considered to be "Ancient". Both Latin and Ancient Greek are still taught in schools to kids who get a classic western education.

According to the definition Latin is ancient and consequently it must have gone extinct more then a millenium ago. However, the Roman Catholic Church has continued to use Latin as its language. It maintains a dictionary of Latin modern vocabulary. Surely Latin may be old but it never went extinct.

Ancient Greek does not qualify as ancient because 1453 means less then a millenium. Ancient Greek is taught in school. Books, like the Harry Potter books are translated in Ancient Greek. As far as I understand it, there has not been a similar usage for Ancient Greek as it existed for Latin.

When you are to tag a text using the ISO-639 codes and its definitions, a modern text in Latin or Ancient Greek cannot be tagged. The first issue is that the definitions clearly limit the time when texts are to be considered in a historical or ancient language. The second issue is that in order to write a modern text neologisms are needed and/or existing words with a modern meaning are needed to express modern concepts.

When the definitions preclude the tagging of the modern expressions of Latin or Ancient Greek, it means that either a new code is needed to indicate the modern expression or the defintions of these languages are wrong.

I would argue that when a language has not seen continued use, the modern text is assigned a separate code. It is distinctly different and by tagging it as such, it may be clear to the reader of a text that the understanding of such a text does not reflect the language and the time when it was a living language. I would argue for a separate ISO-639-3 code.

My question is what do you think about this ?
Thanks,
GerardM

Monday, April 7, 2008

WLDC Conference 2008

The World Language Documentation Centre, together with Bangor University and Language Standards for Global Business, wishes to announce a major multidisciplinary conference to celebrate 2008 as the International Year of Languages

August 22-23, 2008

To be held at the Bangor University Business Management Conference Centre

This event is supported by the Welsh Assembly Government and the UK National Committee to UNESCO

The United Nations announced that 2008 would be the International Year of Languages, recognizing the importance of multilingualism in supporting international understanding. The GUM3C conference will attempt to bridge the communications gap between academia and industry, asking (and attempting to answer) such questions as:


How can industry help academia prioritize its research in the 3 Ms?

What are the developing standards, who are developing them and will they be used?

How will this generate peace, prosperity and global understanding?

Mor info, details on submission of papers or workshops as well as conference registration can be obtained from http://www.gum3c.org

Monday, March 31, 2008

online African dictionaries planning meeting

The Kamusi Project (www.kamusiproject.org) is pleased to announce that
we are about to begin work on PALDO: the Pan-African Living Dictionary
Online. PALDO will build on the Kamusi architecture to create an
interlinked multilingual dictionary for African languages, creating a
powerful communications tool that will be useful throughout the African
continent.

The first step for PALDO will be to program the database, multilingual
tools, and enhanced user interface. This work will begin on April 2
with our partners at Kasahorow (www.kasahorow.org), at a meeting in
Accra, Ghana. This meeting will be simulcast LIVE ONLINE, and the
transcript will also be posted on a special blog at
www.kamusiproject.org/paldo

WE INVITE YOU TO PARTICIPATE IN THE MEETING by logging into the live
chat session or commenting on the blog, beginning at 9 a.m. Ghana time
on April 2. The timezone for Accra is GMT.

We are particularly hoping for participation from:
1) computer programmers and database specialists
2) linguists, lexicographers, and people with an interest in languages
3) users of the Kamusi Project, kasahorow, or other online dictionaries
4) people interested in helping shape the next generation of tools for
African languages

If you would like to participate in this meeting online, please visit
www.kamusiproject.org/paldo for more information.

If you are in Accra and would like to attend in person, the meeting will
be held at the Kofi Annan Centre, beginning at 9 a.m. on Wednesday, April 2.

NOTE: THE DATE OF THE MEETING MAY BE MOVED BACKWARDS ONE DAY IF A 45
MINUTE FLIGHT CONNECTION IS MISSED IN ROME. If the meeting is
rescheduled for April 3, we will place a notice on
www.kamusiproject.org/paldo

Thursday, March 20, 2008

Evolution in the Open Source world

OmegaT is a great open source CAT tool. It is written in Java, it has a growing group of users and Sabine Cretella, who is my weather cock for what is happening in this space, has been a long time champion of the software. As OmegaT makes sense to me for several of the things I am involved in, I have invested and I have been looking for funding to expand its functionality.

Yesterday I was astounded by Sabine. "Anaphraseus", she says, "is a CAT tool that does the things that are critical to me. It allows me to translate into Neapolitian properly; it allows me to enter nap, the ISO-639-3 code and consequently I am able to build my translation memory without having to remember what code I used in stead. They do not have a proper TMX yet, but they are working on it. Now given that I can finally work properly in my language, who cares that I do not have it yet?"

Anaphraseus used to be called "Open Wordfast" and makes use of the Open Office macro tool. It uses the same translation memory format like Wordfast, it supports text segmentation and it is great for proof reading.

Sabine has been investigating Anaphraseus's functionality and so far she is quite pleased. When I asked her why the change, she said that she had been asking for the ISO 639-3 support for almost two years, it was not forthcomming and Anaphaseus is as good for the job.
Thanks,
Gerard