Wall (7,436 threads)
Tips
Before asking a question, make sure to read the FAQ.
We aim to maintain a healthy atmosphere for civilized discussions. Please read our rules against bad behavior.
LeviHighway
10 days ago
jan_OkulaJu
10 days ago
Wezel
13 days ago
odexed
14 days ago
Snark
15 days ago
EugeneGS
15 days ago
Snark
15 days ago
rul
17 days ago
LeviHighway
17 days ago
rul
17 days ago
I think there is a massive problem with the system: The absence of a "pivot" sentence.
In effect, I see numerous cases when a sentence s1 is given in language l1 then translated into s2 in l2, then retranslated from s2 in l1 resulting in s1' and re-retranslated in l3, etc...this is endless, because s1, s1', s1''... are all different...
This happens all the time when there are no one-to-one translation or words don't correspond 100% and the meaning slips slowly from one translation to the next, which happens actually in a majority of the cases because no 2 words ever perfectly match in 2 differnet languages...
There is also an explanation about the structure of the corpus here:
http://blog.tatoeba.org/2010/02...eba.html#rule2
These are really basic explanations, but if you take the time to think about it, you'll find out that the system is very coherent :)
Already we see that direct translations from l1 to l2 and to l3 end up with 2 sentences in l2 and l3 that don't match in most cases, so the discrepancy is just too important.
Subsequently, I think retranslations, that compound theses discrepancies, shouldn't simply appear on the same sentence detail. Retranslations are just different sentences altogether when they happen to make any sense at all...
I think, that as this ramps up, the problem indicated will become more serious. At least it might be worth thinking about whether the system could detect a possible game of "telegram" going on, and warn people (e.g. warn people if they are editing a sentence that is the source of a bunch of sentences in other languages).
As for me, I will just not bother with retranslations anymore, they're just a waste of time!
Yes, what you are talking about is more an issue on the frontend (the user interface). People who are not familiar with the concept and who don't understand the structure are more likely to do things that break the system...
But this is why we have this article:
http://blog.tatoeba.org/2010/02...n-tatoeba.html
It's a difficult project, and it takes a lot of time to come up with an adequate interface. So until we get there, we count on veterans who understand the system to educate new users so they understand how to use the system properly.
In fact there's always a pivot sentence, the one which is displaying first, and yes we all know that translating a translation make each time the meaning slip slowly, that's why a translation of translation is displayed differently (in gray) this way you know that this sentence is not reliable as a "translation" of the current pivot sentence
this is also one of the interesting point of this project, to see how the meaning can slip, as we have a way to tell which sentences are direct translation, and which are not
[en] How can I contribute to translating this web site? I'd like to translate it into Esperanto an Breton.
[fr] Comment puis-je aider à traduire ce site ? J'aimerai le traduire en espéranto et en breton.
A translation branch for Esperanto has recently been created:
https://translations.launchpad..../eo/+translate
Nothing is translated yet, so you can register and translate as much, or as little, as you want.
If that's not enough to keep you busy I'm sure Trang would be happy to add Breton as another branch.
Here for Breton:
https://translations.launchpad..../br/+translate
Thank you very much for offering you help :)
the response time is lagging terribly...
How does one get languages added? I believe I could encourage contributions in Slovenian and in CycL (there's a wikipedia page describing this language ), and I'd love to see contributions in Māori.
Hi!
Suggestions and links showing which flags could be used is really appreciated :)
For Slovenian of course it's Slovenia flag, for Maori I've found this list:
http://en.wikipedia.org/wiki/Li...gs#Maori_flags
This one http://en.wikipedia.org/wiki/Tino_rangatiratanga would be good? (It has an amazing design!)
For CycL, do you have any ideas in particular?
Thank you, and welcome!
I've put in some example CycL sentences, but I had to mark them as English (well, I didn't have to, but marking them as French or Swahili would have been worse). Sentence nº427971 is an example.
I'm not sure about a flag for CycL. It will be an interesting day when AI systems feel the need to have flags. I think it would be OK to use the OpenCyc logo from the top of http://sw.opencyc.org/ if you wanted.
Tatoeba uses 3-letter language codes, and CycL is expected to get a general tag for artificial languages (later it may have a subtag, but suptag support is only planned), so some general symbol for constructed languages may be OK too, IMHO.
By the way, I'm really happy you'll add Slovenian! I've been fascinated by this language for a long time. ^^
@witbrock
Please do this steps http://tatoeba.org/eng/faq for Maori and tell me if the flag I showed you would be good :)
Thank you!
By the way, when you first said flag, I thought you meant "tag" so that you could find them, .... hence the comments on the sentences.
You can recheck, but I tagged them too :)
Any guesses when the CycL flag will be added. I'm eager to relabel my sentences away from English, and to add more!
I think you can go on adding sentences even if the language has not a flag yet :)
This can be a pain for non-trusted users whose new language is autodetected as something else, though...
It will be added some time next week :) Earliest would be August 2nd. Latest would be August 7th.
The CycL flag that you sent looks good! Thank you.
For Maori, I can't actually carry out the last step, since, alas, I don't speak Maori. I'll try to elicit them though!
the basic way to do, is to add some sentences in the language you want, and tell us if there's a particular flag to use, this way we're sure we will not have a 0 sentences languages (because we get a loooot of request of languages), and after in the following week we will add all what is needed.
thanks for your interest in the project :)
It would be great if the number of sentences was dynamically updated. Or even updated whenever the front page loaded. This should be cheap to do.
They are. I just added a fake Danish sentence and the number instantly doubled from one to two.
comment est-on informé qu'une traduction est validée ?
Je viens de créer une phrase et sa traduction et je ne la retrouve pas quand je recherche un mot-clé...
The search index is not updated instantly. In a day or two your sentence will show up in search results.
oui nous allons ajouté un petit texte précisant le fait que l'index est mis à jour une fois par semaine, et que donc les phrases venant d'être ajoutées ne sont pas tout de suite "trouvable" via le moteur de recherche (cela est dû à la petitesse de notre infrastructure)
EN: It would be fine to have a versioning system. This way, we would just update instead of downloading the whole files.
EO: Bonus havi deponejan sistemon. Tiamaniere, oni nur ĝisdatigi ol elŝuti tutan dosieron.
Linux / Unix users might be well served by use of an rsync server. I never got on very well with it when I tried to use one with Windows.
EN: You're right, but a versionning system would be more effecient as it'd only send differences between states. And we would have all the past versions.
EO: Vi pravas, sed deponeja sistemo estu pli efika pro tio, ke ties elsendo estus nur malsamoj inter statoj. Kaj oni havus ĉiujn pasintajn versionojn.
+1 I've never tought about it, but it's for sure a good idea.
How come everything here is in English in this "multilingual" project ?!?
not everything is in english, and you could have posted in your mother tongue, we obliged nothing, if you pay attention, comment on french sentences are most of the time in french, comment on chinese sentences are in chinese etc.
it's up to everyone to choose the language he wants to communicate here,
if the interface is not complete or not present in your language, you can ask help, and we will be glad to show you how to help us translating tatoeba in your languages.
you should start with removing the english word 'Project' from your logo...That would help!
Then we could remove the Japanese word Tatoeba from the logo ... :-P
a name and a descriptive word are quite different entities..."Project" definitely qualifies your project as an English-language one. It thus uselessly deters non-English speakers, creating a bias in your contributors representation. Is that what you're seeking ?
For all that this thread is a bit tendentious, on purely aesthetic grounds, not having "project" in the logo would probably be an improvement. I think that calling the project "tatoeba" is probably better too (e.g. Wikipedia is just called "Wikipedia" rather than "Wikipedia Project").
'everything' isn't.
First, if you go to the top right hand corner you can select eight other languages for the interface.
Second, comments on examples are often in the language of those examples.
Third, Sbgodin's post just below yours is in Esperanto as well as English.
Il n'empêche, la grande majorité de ce que je vois ici est en anglais et les quelques phrases en français sont incorrectes...
Exemple, dans la rubrique "Astuces": "Ici vous pouvez DEMANDER des questions..."
En français, on ne "demande" pas des questions, on les "'pose".
les quelques (plus de 30 000, 3ème langue en terme de phrases et surement première en terme de nombre de contributeurs) phrases française proviennent en majorité d'un autre projet et sont, en effet pour la plupart, erronées
tout cela, ainsi que le moyen de faire la différence entre ces phrases, et les phrases ajoutés par les contributeurs sont expliqués dans cette article
http://linuxfr.org/2010/07/17/27136.html
un des buts de tatoeba est donc de corriger cela
pour l'interface en elle même, de nombreux pans ont été rajoutés il y a peu, et nous n'avions pas encore eu le temps de la traduire en français.
pour l'erreur, en effet je te remercie de nous le faire remarquer, je l'ai corrigé à l'instant et à la prochaine mise à jour (qui se fera ce weekend) la faute aura disparue,
il faut garder en tête que le projet est collaboratif, et qu'il tient à chacun, de nous aider à en améliorer le contenu, et corriger petit à petit les erreurs.
Pour ce qui est de la prédominance de l'anglais (qui est loin d'être le cas dans les récents ajouts de phrases), je ne peux que t'encourager à nous aider à réunir une communauté de francophone.
si tu vois d'autres erreurs dans la version française, c'est avec joie que je les corrigerais
merci :)
OK. Je souhaite participer. Je peux contribuer en français/anglais/allemand/néerlandais.
Qu'est-ce que l'UTC ?
très bien :)
UTC = Université de Technologie de Compiègne (c'est une école d'ingénieur française, dont fait partie Trang, la fondatrice et guru de tatoeba)
est-il possible d'importer des phrases "en masse" ?
Je dispose d'environ 4000 phrases français/allemand
Que se passe-t-il si une même phrase dans une même langue est créée plusieurs fois par plusieurs contributeurs lors de traductions à partir de langues différentes ? Est-ce que Tatoeba opère le rapprochement automatique ?
les modérateurs peuvent importer des phrases en masses,
donc tu peux nous envoyer ta liste à team@tatoeba.fr
cette fonctionnalité n'est pas disponible pour tout le monde pour l'instant, pour éviter que des petits malins ajoute des masses de phrases, ce qui aurait pour conséquence de ralentir terriblement le site, de plus nous devons vérifier que les phrases de la liste ne sont pas des copier coller de livres ou autres médias soumis aux droits d'auteurs (car vu que nous redistribuons sous licence libre l'ensemble de la phrase, ce serait caduque de mettre dans tatoeba des phrases provenant de manuel scolaire / de dictionnaire / de livres pas encore dans le domaine public)
il y a bot de détection des doublons qui tourne une fois par semaine (environ) et qui fusionne les traductions identique dans la même langue, comme ça les utilisateurs peuvent traduire sans se soucier de cela.
et que se passe-t-il quand le robot détecte un conflit de traductions ?
Je comprends qu'un contrôle de l'apport en masse est nécessaire, bien sûr.
Dans quel format voulez-vous les traductions ?
la phrase la plus ancienne possédant un propriétaire est gardé, les autres doublons supprimés, et les traductions/commentaires/listes/tags/audios associés au phrases supprimés sont réassociés à la phrase qui est gardé
le format est
phrase1[tabulation]traduction1
en gardant le même ordre (toujours français en premier, ou toujours allemand en premier) si tu les as sous un ordre format, (phrase1[virgule]traduction1) je peux me charger de le convertir.
Suggestion: Import list function
This is a bit of a minority interest suggestion, but I think is should be quite easy to implement. It's possible to download lists, but it's not possible to 'upload' lists. What I suggest is, next to the "Create a new list" option, a new 'Import list' function.
Basic idea - upload (or paste into a multi-line text box) a list of sentence numbers separated by new lines. Result - a list including those sentences.
Fancy variants - Same thing, only using a list of the sentence text instead works as well.
modos can already do this
the reason why it's not possible for everyone is
It's a bit slow yet, moreover this will cause some problem
what if I download a list, correct it, and reexport it, it would even slower to check if sentences come from a previous export, what if I've corrected a lot of sentences but people has also corrected sentences directly in tatoeba.
moreover it will be even harder to check if people add appropriate content if they add bunch of hundred of sentences.
but as said, you as a modos, and people, by sending us or to you, the list, can import list of sentences
> modos can already do this
Nope. You've misunderstood what I meant.
I'm not talking about importing _sentences_. I'm talking about importing a _list_. That is, the end result is a new list of existing sentences.
ok I see
basically you will have a file with a list of id and it will create the list with all these ids being part of this list ?
That's it.
Could be very useful to me - or anybody who uses the download files.
as ever if you give us such a list, we can do it, it can be a temporary solution
Well there is one I could do with - but it's only 41 entries so it doesn't seem worth getting you to do it manually.
Trusted user nomination.
I'd like to suggest qahwa as a suitable candidate for 'Trusted User' status. We could certainly use a few more native speakers of Japanese active in Tatoeba.
Okay, I sent qahwa a private message about this :)