Wall (7,463 threads)
Tips
Before asking a question, make sure to read the FAQ.
We aim to maintain a healthy atmosphere for civilized discussions. Please read our rules against bad behavior.
Tamajeq1286
7 hours ago
Thanuir
yesterday
LeviHighway
2 days ago
AlanF_US
4 days ago
leucammieuvu
4 days ago
LeviHighway
5 days ago
AlanF_US
6 days ago
araneo
8 days ago
araneo
8 days ago
LeviHighway
8 days ago
Following up on my comment here: http://tatoeba.org/eng/sentences/show/477093
I wanted to know what our general stance is on adding machine translations?
Thanks to notice it up, I think we need to have a discussion with him. For adding machine translations, we do not want this. The goal of tatoeba is precisely to propose a better option for people willing to have a set of sentences translated in an other language,
Good to know. Thank you everyone for your replies!
I don't think it's ever come up before - but, bad, bad idea.
> 4. Do not translate word for word
>
> We are not interested in having sentences that sound like
> they were written by a robot. We want sentences that really
> are what a native speaker would say.
Tatoeba Blog, ‘How to be a good contributor’, rule 4 (http://blog.tatoeba.org/2010/02...eba.html#rule4)
I guess this covers machine translations. ^^
Hi,
should I add colloquialisms as translations?
Thanks
I don't see why not, just indicate that the translations are colloquial in a comment and a trusted user will add a tag to it...tatoeba already has colloquial sentences http://tatoeba.org/eng/tags/sho...tag/Colloquial
or read this http://blog.tatoeba.org/2010/02...n-tatoeba.html and then PM Trang (http://tatoeba.org/eng/private_...es/write/TRANG) and ask her nicely to make you a trusted user...then you can add your own tags
you can read more about tatoeba's submission policy here:
http://blog.tatoeba.org/2010/08...f-content.html
hope I helped
Yes, you did! Thanks guys!
You can. The only rule is to have the best compromise between "be close to the meaning" and "being natural". If it's a formal sentence, add a translation which is also formal etC... If you think two translations can be added, you can add both, and maybe add in comment the specificity of each sentences (if one is more natural but not so close to the original meaning, and the other closer to the original meaning but maybe less natural for example).
Je vois de nombreux doublons au point près, c'est à dire que la seule différence est la présence ou non du point final.
La procédure de déduplication ne les considère-t-elle pas comme des doublons ?
Non le script ne détecte que les doublons parfait, je n'ai pas préféré ajouté cela, car le script de détection est assez "sale" (c'est une procédure en pl/sql de mysql) et je n'ai pas trouvé de manière propre et tout aussi ""rapide"" de le faire. De plus dans certains cas la ponctuation peut changer la traduction (même si je te l'accorde l'absence de point final ne change en rien le sens vu que c'est purement et simplement une faute).
Mais je pense qu'il faudrait pour être plus "propre" créer un script d'uniformisation, correction des majuscules, des points manquants etc. qui tourneraient avant le script de dédoublonnage. Cependant je ne pense pas avoir le temps ces prochaines semaines, de faire une telle chose.
Par contre en attendant je pense pouvoir créer petit à petit un jeu de commande sql de "nettoyage" pour ajouter les points manquants etc.
Merci pour ta réponse. Vivement l'ajout automatique de points finaux, ça va résoudre plein de problèmes !
Mais je ne sasi pas si tu avais vu mon message précédent concernant les phrases nouvelles - donc non encore dédupliquées - qui apparaissaient dans les résultats de recherche.
En tout état de cause, à quelle fréquence fais-tu passer la procédure de dédoublonnage ? Le savoir me permettrait d'éviter de traduire des doublons potentiels en regardant leur date/heure de création.
Ce qui est dommage, c'est que ce sont toujours ces dernières phrases, souvent parasites, qui apparaissent toujours en premier dans les recherches...
Il est normalement lancé une fois par semaine. il a été lancé il y a deux semaines, mais visiblement une information pour l'indexation avec un autre projet était "écrasé" parce script. Donc je dois (encore) le remodifier avant de le relancer.
Sinon je peux changer le sens d'affichage des résultats de recherches, et afficher les plus anciennes d'abord ?
Je ne sais pas si il faut d'abord afficher les plus anciennes (parce qu'il y a plein de vieux tromblons...) mais il faudrait peut-être "daplacer" les 10 derniers jours (ou les phrases créées depuis la dernière procédure de déduplication) à la fin...
Voir ce qu'en pensent d'autres traducteurs en masse...
*"déplacer"
Tatoeba beta testing
Do you get to play with the real data set, or only a 'safe' copy?
If you're talking about the "blue" version of tatoeba, it's a separate database with different password , so changing one does not change the other.
Can someone confirm this problem I'm having? I created a few tags yesterday:
- cmt: alternative vocabulary,
- cmt: alternative grammatical number, and
- cmt: alternative gender
by adding them to
* http://tatoeba.org/eng/sentences/show/478530, and
* http://tatoeba.org/eng/sentences/show/478533
They appear at the bottom of
* http://tatoeba.org/eng/tags/view_all
but the pages to which the tags link look like the tag doesn't exist, e.g.
* http://tatoeba.org/eng/tags/sho...rnative_gender
These are the only tags with a colon, so I suspect that's the problem. I tried escaping the colon in the URL,
* http://tatoeba.org/eng/tags/sho...rnative_gender
but that doesn't work. Is this a known problem? Should I file a bug report somewhere?
Yes indeed. The colon is the problem. Thanks for reporting. I created a ticket for it.
Thanks! Looking forward to the update, by the way.
ow to use totoeba
by asking here :)
welcome!
(it's tAtoeba ^^)
If you're looking for example illustrating a word you can simply use the top menu bar in green, for example you want sentence using "example" in English sentences
http://tatoeba.org/eng/sentence...&query=example
If you see a translation is missing in your native language, you simply add translation by clicking on the "translate button" http://flags.tatoeba.org/img/translate.png
If you want to add a really new sentences you can do this by going on the contribute page
http://tatoeba.org/eng/contribute
If you have a specific need, ask, maybe it's already possible in tatoeba, or at least we can think about making it possible :)
Tell us if you have questions
Je pense que les phrases qui viennent d'être créées et ne sont pas encore dédupliquées ne devraient pas être disponibles dans les résultats de recherche, parce qu'alors on exacerbe le problème en créant de nouvelles traductions de copies.
More on tags...
I trimmed the deleted tags from my list and added the number of sentences tagged (according to the latest tags.csv file).
There are still a few that need to be deleted and I've listed these at
http://martin.swift.is/tatoeba/tags.html#delete
Then there's the bunch that needs to be renamed. Can I just send someone a tab-separated file with the old and new names? Most are simple transliterations, but some will need to be discussed.
For example: there are several "check" tags. I suggest we merge "to check", "@Needs Native Check", "Needs Native Check" and "@grammar check" under "@check", "@native check" or something similar ("to be reviewed" as well?).
Then "delete" should be merged with "@delete". I take it this would be considerably simpler with direct database access than through the interface (though I guess one could cURL one's way through that).
[not needed anymore- removed by CK]
Sorry, I failed to mention that I didn't really clean my parsing code up that well; didn't bother with false positives (I'll look into passing grep a more nuanced pattern next time around) and the script doesn't handle tags with parentheses.
I fixed the "British" tag by hand in case you want to resort your page, CK.
PG-13 is a normal tag.
be-1959acad shows that this is applicable only in the Academic variery of Belarusian; it's a normal language tag
Leopolis is a latin duplicate of Lviv ;o
Urdu is to be deleted, since we have a flag now
IMHO spoken to male should be replaced with ‘said to male’
Thanks, Demetrius!
I've moved be-1959acad to the "Language" section, Leopolis to "redundant" and Urdu to "depopulate".
Regarding PG-13, I'm just thinking that a more culture independent tag would be more useful.
Since there is only one "spoken to male" tagged sentence and an existing "said to male" tag, I'll just re-tag that sentence and move the tag to "empty".
Azeri is a language spoken in Azerbaijan.
Thanks!
By the way, about the XXX tag...
How it should be used? Consider the following groups:
a) sentence describes a sexual intercourse in rude words,
b) sentence uses the rude words with an indirect meaning, to describe something other,
c) sentence describes a sexual intercourse with euphemisms
In which cases XXX tag should be used?
I would say a) and c) and possibly b) (but depending on circumstances).
We'll probably need to fine tune that sort of thing later. It's not really a priority at the moment as the tags don't actually do much yet.
I don't think these should should be tagged under a single label. I think a) and b) could be tagged with something like "rude" or "obscene" and a) and c) with something like "sex" or "pornographic". c) should furthermore be tagged with "euphemism".
I this is actually a great time to think about how we're using and how we'd like to use the tags, seeing how it's not a priority and we can play around with them.
Should we keep proverbs offending other nations in the database? And how should these be tagged?
We already have some Ukrainian proverbs about Russians in the database (sth like “You can ward off the devil crossing yourself, but you can’t ward off a Moskal”). ^^
You could tag them as "stereotype", "prejudice" or even "bigotry".
I'm all for freedom of expression, but I think that clear lines are more useful for a project such as this (oh, and trying to steer around hypocritical stances would be nice). Seeing how the aim of this project is to gather sentences and link them to translations -- not to disseminate facts -- and the web is open to anyone wanting to espouse their hate, I'm not overly concerned about losing offensive content.
I do, however, prefer the idea of simply filtering out the filth one doesn't like.
There are already quite a few labelled as 'Lie'. Filters would be nice (filtering out XXX sentences should be a priority if we want to be 'school friendly'), but I think that sort of content is still needed (in moderation) to give a full coverage of language usage.
> full coverage of language usage.
Talking about which, I'm reminded of a Japanese speaker who was certain that a certain body part was referred to as "pussy cat". Nothing I could say would persuade him that the 'cat' wasn't needed; he just said "I use it all the time with my girlfriend so it must be right."