[32319] in Perl-Users-Digest
Perl-Users Digest, Issue: 3586 Volume: 11
daemon@ATHENA.MIT.EDU (Perl-Users Digest)
Tue Jan 10 18:09:27 2012
Date: Tue, 10 Jan 2012 15:09:08 -0800 (PST)
From: Perl-Users Digest <Perl-Users-Request@ruby.OCE.ORST.EDU>
To: Perl-Users@ruby.OCE.ORST.EDU (Perl-Users Digest)
Perl-Users Digest Tue, 10 Jan 2012 Volume: 11 Number: 3586
Today's topics:
Language of the year <cartercc@gmail.com>
Re: Language of the year <bugbear@trim_papermule.co.uk_trim>
Re: POD module synopses? <nospam-abuse@ilyaz.org>
Re: POD module synopses? (Seymour J.)
Re: POD module synopses? <nospam-abuse@ilyaz.org>
Re: POD module synopses? <nospam-abuse@ilyaz.org>
Re: regexp for matching a string with mandatory undersc <nospam-abuse@ilyaz.org>
Re: regexp for matching a string with mandatory undersc <ben@morrow.me.uk>
Re: regexp for matching a string with mandatory undersc <rweikusat@mssgmbh.com>
Re: regexp for matching a string with mandatory undersc <hjp-usenet2@hjp.at>
Re: regexp for matching a string with mandatory undersc <nospam-abuse@ilyaz.org>
Re: regexp for matching a string with mandatory undersc <ben@morrow.me.uk>
Re: regexp for matching a string with mandatory undersc <ben@morrow.me.uk>
Re: regexp for matching a string with mandatory undersc <rweikusat@mssgmbh.com>
Re: UCD, and Unicode-AGE of a character? <nospam-abuse@ilyaz.org>
Re: UCD, and Unicode-AGE of a character? <ben@morrow.me.uk>
Unicode-AGE of a character? <nospam-abuse@ilyaz.org>
Re: Unicode-AGE of a character? <ben@morrow.me.uk>
Digest Administrivia (Last modified: 6 Apr 01) (Perl-Users-Digest Admin)
----------------------------------------------------------------------
Date: Mon, 9 Jan 2012 09:21:16 -0800 (PST)
From: ccc31807 <cartercc@gmail.com>
Subject: Language of the year
Message-Id: <f8f44c5d-df3f-49e6-8a76-3e0166965cdc@v14g2000yqh.googlegroups.com>
For those of you who follow these things, and with apologies to those
of you who don't, the TIOBE index has awarded its Language of the Year
designation to Objective-C.
C# has replaced C++ as the 3rd place language, Perl maintains its 9th
place status, and Lisp its 13th place. PHP, Python, and Ruby continue
to drop.
Yeah, I know that this 'information' is more or less meaningless, but
I find this kind of rating interesting even if mostly meaningless.
http://www.tiobe.com/index.php/content/paperinfo/tpci/index.html
All the best for the new year, CC.
------------------------------
Date: Tue, 10 Jan 2012 09:19:47 +0000
From: bugbear <bugbear@trim_papermule.co.uk_trim>
Subject: Re: Language of the year
Message-Id: <-sOdnRfbLfYun5HSnZ2dnUVZ8lKdnZ2d@brightview.co.uk>
ccc31807 wrote:
>
> Yeah, I know that this 'information' is more or less meaningless,
Thanks for posting.
BugBear
------------------------------
Date: Tue, 10 Jan 2012 06:59:42 +0000 (UTC)
From: Ilya Zakharevich <nospam-abuse@ilyaz.org>
Subject: Re: POD module synopses?
Message-Id: <slrnjgnoeu.4ik.nospam-abuse@panda.math.berkeley.edu>
On 2012-01-03, Ben Morrow <ben@morrow.me.uk> wrote:
> Unless you have a particular need for IPF documentation, I'd recommend
> using something else. If you don't find command-line perldoc acceptable,
> you can build an HTML tree with something like
I doubt that anyone who used IPF version for a couple of hours would
be ever satisfied with what perldoc or pod2html does... (However, I
did not try running the stuff the way you recommended, so my opinion
here is based on PRIOR experiences...)
Ilya
------------------------------
Date: Tue, 10 Jan 2012 10:23:23 -0500
From: Shmuel (Seymour J.) Metz <spamtrap@library.lspace.org.invalid>
Subject: Re: POD module synopses?
Message-Id: <4f0c57eb$1$fuzhry+tra$mr2ice@news.patriot.net>
In <slrnjgnoeu.4ik.nospam-abuse@panda.math.berkeley.edu>, on
01/10/2012
at 06:59 AM, Ilya Zakharevich <nospam-abuse@ilyaz.org> said:
>I doubt that anyone who used IPF version for a couple of hours would
>be ever satisfied with what perldoc or pod2html does...
Would you be willing to look at my attempts to correct the code and
try to spot what I'm missing? If so, would you want a diff for what
I've tried or the modified code itself?
Also, do you know anything about the provenance of the 1.34 version?
The author is listed as "vera", although there is no vera in the
AUTHOR section of the POD. Or is vera an alias for you?
--
Shmuel (Seymour J.) Metz, SysProg and JOAT <http://patriot.net/~shmuel>
Unsolicited bulk E-mail subject to legal action. I reserve the
right to publicly post or ridicule any abusive E-mail. Reply to
domain Patriot dot net user shmuel+news to contact me. Do not
reply to spamtrap@library.lspace.org
------------------------------
Date: Tue, 10 Jan 2012 19:36:03 +0000 (UTC)
From: Ilya Zakharevich <nospam-abuse@ilyaz.org>
Subject: Re: POD module synopses?
Message-Id: <slrnjgp4p3.6kj.nospam-abuse@panda.math.berkeley.edu>
On 2012-01-10, Shmuel Metz <spamtrap@library.lspace.org.invalid> wrote:
> In <slrnjgnoeu.4ik.nospam-abuse@panda.math.berkeley.edu>, on
> 01/10/2012
> at 06:59 AM, Ilya Zakharevich <nospam-abuse@ilyaz.org> said:
>
>>I doubt that anyone who used IPF version for a couple of hours would
>>be ever satisfied with what perldoc or pod2html does...
>
> Would you be willing to look at my attempts to correct the code and
> try to spot what I'm missing? If so, would you want a diff for what
> I've tried or the modified code itself?
My development machine would not go out of BIOS, and I did not prepare
for this - not enough "live boot in from CD or emulator" images, and
no images to boot on newer hardware. (So I cannot even read files on
the non-backed-up partitions - including my most recent builds of
Perl/OS2. I have all the sources backed up, but did not do inventory
of which patches were backed up from, and "apply in which order"...
Sigh... It is a question of time - I know HOW to recover, but all the
scenarios I have in mind are so convoluted that I hesitate to start.)
> Also, do you know anything about the provenance of the 1.34 version?
> The author is listed as "vera", although there is no vera in the
> AUTHOR section of the POD. Or is vera an alias for you?
"vera" is just what rcs puts in on my development machine.
Ilya
------------------------------
Date: Tue, 10 Jan 2012 19:46:33 +0000 (UTC)
From: Ilya Zakharevich <nospam-abuse@ilyaz.org>
Subject: Re: POD module synopses?
Message-Id: <slrnjgp5cp.6kj.nospam-abuse@panda.math.berkeley.edu>
On 2012-01-05, Shmuel Metz <spamtrap@library.lspace.org.invalid> wrote:
> Some things that are definitely bugs in pod2ipf are
OK, OK, I run through backups, and one nearby has the rcs file. So
the version in ilyaz.org/software/os2/pod2ipf.zip may have a wrong
timestamp, but should be otherwise identical to my latest and greatest
version.
Enjoy,
Ilya
------------------------------
Date: Tue, 10 Jan 2012 06:44:50 +0000 (UTC)
From: Ilya Zakharevich <nospam-abuse@ilyaz.org>
Subject: Re: regexp for matching a string with mandatory underscores
Message-Id: <slrnjgnnj2.4ik.nospam-abuse@panda.math.berkeley.edu>
On 2012-01-03, Rainer Weikusat <rweikusat@mssgmbh.com> wrote:
>>>> For which value of "well"? If it is applied to 2GB string, would
>>>> it make a copy of it?
>>>
>>>Not when counting or replacing character in a non-UTF8 string.
So I read it as: "it will" (with certain exceptions).
>>>> If the string is tied to a database entry, would
>>>> it cause a database update?
>>>
>>>Maybe, maybe not. That would depend on the implemention of tieing
>>>mechanism.
Again...
> OTOH, it is sensible to assume that - usually - the people who wrote
> the implementation will have tried to make it behave sensibly and in
> this case, that tr/// will neither copy nor modify the string except
> if this is necessary to perform the requested operation.
Not applicable to Perl (in general). A lot of stuff is majorly pessimized.
> Re: tied scalars
>
> What will happen when an operation is performed on a scalar tied to
> something depends on the class/ module used to provide the tied
> semantics and this can be anything, so the question didn't really make
> sense: This class or module may well cause 'a database update' despite
> perl didn't modify the data.
This is true "literally", but AFAIK, not applicable to any situation I
know.
Essentially, for me all this boils down to: do not use tr/// unless
you can't avoid it, or know EXACTLY how and when your code is going to
be used...
Ilya
------------------------------
Date: Tue, 10 Jan 2012 11:36:05 +0000
From: Ben Morrow <ben@morrow.me.uk>
Subject: Re: regexp for matching a string with mandatory underscores
Message-Id: <5ektt8-c3m.ln1@anubis.morrow.me.uk>
Quoth Ilya Zakharevich <nospam-abuse@ilyaz.org>:
> On 2012-01-03, Rainer Weikusat <rweikusat@mssgmbh.com> wrote:
>
> > OTOH, it is sensible to assume that - usually - the people who wrote
> > the implementation will have tried to make it behave sensibly and in
> > this case, that tr/// will neither copy nor modify the string except
> > if this is necessary to perform the requested operation.
>
> Not applicable to Perl (in general). A lot of stuff is majorly pessimized.
RTFS. tr/// is pretty highly optimised; in particular, the 'count
characters' case has its own implementation that does no copying, except
when counting non-SvUTF8 characters in a SvUTF8 string. In that case
obviously every character of the string being counted has to be
individually converted to UTF-32, so there's no allocation but there is
effectively copying.
This sort of inefficiency is unavoidable when using UTF-8 as an internal
representation, which is why certain people are trying so hard to make
perl's internal representation opaque. Everyone now knows that using
UTF-8 was a mistake, but it can't be fixed until people get used to
keeping their fingers out.
> > Re: tied scalars
In the tied (or more generally magic) case, perl calls FETCH to update
the string stored in the scalar, does the tr/// on that string, then
calls STORE to update the magic. A tie implementation that's being
careful about copying will have no additional problems because of tr///.
Ben
------------------------------
Date: Tue, 10 Jan 2012 17:51:06 +0000
From: Rainer Weikusat <rweikusat@mssgmbh.com>
Subject: Re: regexp for matching a string with mandatory underscores
Message-Id: <874nw3zbsl.fsf@sapphire.mobileactivedefense.com>
Ben Morrow <ben@morrow.me.uk> writes:
[...]
> Everyone now knows that using UTF-8 was a mistake,
That's not something "everyone knows" and in fact, some people were so
convinced that UTF-8 would be a sensible choice that they implemented
complete operating systems based on using UTF-8 as native character
encoding (that would be "Plan9"). This should rather be "every member
of some small group of people" (people currently working on Perl
Unicode support?) are strongly convinced that chosing UTF-8 was a
mistake (and I'd wager a bet that the base reason for this is "that's
not what Microsoft did and consequently, it must be WRONG !!1").
------------------------------
Date: Tue, 10 Jan 2012 20:11:39 +0100
From: "Peter J. Holzer" <hjp-usenet2@hjp.at>
Subject: Re: regexp for matching a string with mandatory underscores
Message-Id: <slrnjgp3bb.v35.hjp-usenet2@hrunkner.hjp.at>
On 2012-01-10 11:36, Ben Morrow <ben@morrow.me.uk> wrote:
> RTFS. tr/// is pretty highly optimised; in particular, the 'count
> characters' case has its own implementation that does no copying, except
> when counting non-SvUTF8 characters in a SvUTF8 string. In that case
> obviously every character of the string being counted has to be
> individually converted to UTF-32, so there's no allocation but there is
> effectively copying.
>
> This sort of inefficiency is unavoidable when using UTF-8 as an internal
> representation, which is why certain people are trying so hard to make
> perl's internal representation opaque. Everyone now knows that using
> UTF-8 was a mistake, but it can't be fixed until people get used to
> keeping their fingers out.
In Pike (like Perl a vaguely C-like interpreted language) strings always
consist of elements of equal length: All characters in a string
are either 1 byte or 2 bytes or 4 bytes in length. That may waste some
space if you have a string with lots of ascii characters and one 💩 in
it, but it makes most string operations simpler.
Theoretically, Perl could switch to such a model without breaking
programs (except XS code). Practically ...
hp
--
_ | Peter J. Holzer | Deprecating human carelessness and
|_|_) | Sysadmin WSR | ignorance has no successful track record.
| | | hjp@hjp.at |
__/ | http://www.hjp.at/ | -- Bill Code on asrg@irtf.org
------------------------------
Date: Tue, 10 Jan 2012 19:56:10 +0000 (UTC)
From: Ilya Zakharevich <nospam-abuse@ilyaz.org>
Subject: Re: regexp for matching a string with mandatory underscores
Message-Id: <slrnjgp5up.6kj.nospam-abuse@panda.math.berkeley.edu>
On 2012-01-10, Ben Morrow <ben@morrow.me.uk> wrote:
>> Not applicable to Perl (in general). A lot of stuff is majorly pessimized.
>
> RTFS. tr/// is pretty highly optimised; in particular, the 'count
> characters' case has its own implementation that does no copying, except
> when counting non-SvUTF8 characters in a SvUTF8 string. In that case
> obviously every character of the string being counted has to be
> individually converted to UTF-32, so there's no allocation but there is
> effectively copying.
What makes this "obvious"? I see absolutely no need for this...
Unless you mean "copying one char at a time", not copying the whole
string. And such things MUST be documented (since in presence of
tie()ing they are not implementation details).
> This sort of inefficiency is unavoidable when using UTF-8 as an internal
> representation, which is why certain people are trying so hard to make
> perl's internal representation opaque. Everyone now knows that using
> UTF-8 was a mistake, but it can't be fixed until people get used to
> keeping their fingers out.
Why do you think it is inefficiency? Todays machines are even more
tied by memory than machines 10 years ago... (In proportion to amount
of data one may [so does] store on the disk.)
>> > Re: tied scalars
> In the tied (or more generally magic) case, perl calls FETCH to update
> the string stored in the scalar, does the tr/// on that string, then
> calls STORE to update the magic. A tie implementation that's being
> careful about copying will have no additional problems because of tr///.
Now I'm absolutely confused... Are you still discussing tr/foo//
here? Do you say it WOULD call STORE?
And "being careful about copying" brings no imagery here. What
EXACTLY do you mean by that?
IMO, an operation which has semantic of reading should NOT call STORE
on tied data...
Ilya
------------------------------
Date: Tue, 10 Jan 2012 20:49:12 +0000
From: Ben Morrow <ben@morrow.me.uk>
Subject: Re: regexp for matching a string with mandatory underscores
Message-Id: <8rkut8-trs.ln1@anubis.morrow.me.uk>
Quoth "Peter J. Holzer" <hjp-usenet2@hjp.at>:
>
> In Pike (like Perl a vaguely C-like interpreted language) strings always
> consist of elements of equal length: All characters in a string
> are either 1 byte or 2 bytes or 4 bytes in length. That may waste some
> space if you have a string with lots of ascii characters and one 💩 in
> it, but it makes most string operations simpler.
Just so. If you're feeling clever you can even make your strings into
lists of sections, where different sections can have different
representations. I believe that's the model Perl 6 uses (though it may
have changed: I haven't been keeping track). Quite apart from the
encoding issue, this lets you do quite a lot of manipulation of a large
read-only string without actually copying anything.
> Theoretically, Perl could switch to such a model without breaking
> programs (except XS code). Practically ...
Yeah, I suspect we'll never get there. Ah well...
Ben
------------------------------
Date: Tue, 10 Jan 2012 21:11:47 +0000
From: Ben Morrow <ben@morrow.me.uk>
Subject: Re: regexp for matching a string with mandatory underscores
Message-Id: <j5mut8-sct.ln1@anubis.morrow.me.uk>
Quoth Ilya Zakharevich <nospam-abuse@ilyaz.org>:
> On 2012-01-10, Ben Morrow <ben@morrow.me.uk> wrote:
> >> Not applicable to Perl (in general). A lot of stuff is majorly pessimized.
> >
> > RTFS. tr/// is pretty highly optimised; in particular, the 'count
> > characters' case has its own implementation that does no copying, except
> > when counting non-SvUTF8 characters in a SvUTF8 string. In that case
> > obviously every character of the string being counted has to be
> > individually converted to UTF-32, so there's no allocation but there is
> > effectively copying.
>
> What makes this "obvious"? I see absolutely no need for this...
> Unless you mean "copying one char at a time", not copying the whole
> string.
Yes, I mean one char at a time. There isn't ever a complete UTF-32 copy
made anywhere, but the equivalent amount of work is done (minus the
allocation).
> And such things MUST be documented (since in presence of
> tie()ing they are not implementation details).
Um, really? What difference does it make? All a tie class needs to know
is that FETCH needs to update the SvPV, and STORE needs to read it out
again. You haven't ever been able to assume perl won't re-allocate the
SvPV when you're not looking, so you have to account for that anyway.
(And yes, this is ridiculously inefficient, but not in any way specific
to tr///.)
> > This sort of inefficiency is unavoidable when using UTF-8 as an internal
> > representation, which is why certain people are trying so hard to make
> > perl's internal representation opaque. Everyone now knows that using
> > UTF-8 was a mistake, but it can't be fixed until people get used to
> > keeping their fingers out.
>
> Why do you think it is inefficiency? Todays machines are even more
> tied by memory than machines 10 years ago... (In proportion to amount
> of data one may [so does] store on the disk.)
You may be right there, I don't know. Using UTF-8 certainly makes the
code a lot hairier, and I suspect that costs more than the memory. You
end up converting character-at-a-time to UTF-32 practically every time
you do anything with that string, rather than being able to use fast
interfaces like wmemchr(3).
> >> > Re: tied scalars
>
> > In the tied (or more generally magic) case, perl calls FETCH to update
> > the string stored in the scalar, does the tr/// on that string, then
> > calls STORE to update the magic. A tie implementation that's being
> > careful about copying will have no additional problems because of tr///.
>
> Now I'm absolutely confused... Are you still discussing tr/foo//
> here? Do you say it WOULD call STORE?
No, I was discussing tr/// in general. That's a good point, though, and
I don't know the answer...
~% perl -E'
sub TIESCALAR { bless [] }
sub FETCH { "foo" }
sub STORE { warn "STORE" }
tie my $x, "main";
say $x =~ tr/o//;
say $x =~ tr/o/o/;
say $x =~ tr/oo/oo/;
say $x =~ tr/oo/o/;'
2
2
2
STORE at -e line 4.
2
So it doesn't call STORE unless you defeat the optimiser by being
excessively cute. I don't know if this should be considered documented,
though I suspect if p5p broke it and someone complained it would get
changed back.
> And "being careful about copying" brings no imagery here. What
> EXACTLY do you mean by that?
Nothing specific, just an implementation (probably in XS) that has been
carefully written not to do any more copying than is necessary. I don't
believe using tr/// will ever lead you to end up making more copies of a
string than you would have otherwise.
Ben
------------------------------
Date: Tue, 10 Jan 2012 21:49:16 +0000
From: Rainer Weikusat <rweikusat@mssgmbh.com>
Subject: Re: regexp for matching a string with mandatory underscores
Message-Id: <87lipfxm77.fsf@sapphire.mobileactivedefense.com>
Ben Morrow <ben@morrow.me.uk> writes:
[...]
> Using UTF-8 certainly makes the code a lot hairier, and I suspect
> that costs more than the memory. You end up converting
> character-at-a-time to UTF-32 practically every time you do anything
> with that string, rather than being able to use fast
> interfaces like wmemchr(3).
Sometimes, life is just mean. Couldn't the people who invented UTF-8
in 1993 for use on their incredibly fast machines have foreseen how
much slower hardware was going to become in the next 19 years?
------------------------------
Date: Tue, 10 Jan 2012 19:28:38 +0000 (UTC)
From: Ilya Zakharevich <nospam-abuse@ilyaz.org>
Subject: Re: UCD, and Unicode-AGE of a character?
Message-Id: <slrnjgp4b6.6kj.nospam-abuse@panda.math.berkeley.edu>
On 2012-01-10, Ben Morrow <ben@morrow.me.uk> wrote:
>
> Quoth Ilya Zakharevich <nospam-abuse@ilyaz.org>:
>> I looked through the docs I could find, and can't find any way to
>> determine the "Unicode AGE" of a particular codepoint except for:
>>
>> a) running /\p{Present_in: FOO}/ for all forseeable values of FOO;
>>
>> b) manually parsing $out = do 'unicore/To/Age.pl';.
>>
>> Do I miss anything?
>
> I don't think so. Note that before (I think) 5.14 unicore/To/Age.pl
> doesn't exist, and before (I think) 5.12 unicode/DAge.txt doesn't exist
> either. You may be better off just grabbing a copy of DerivedAge.txt
> from the Unicode Consortium directly, and using that.
What would be the best fix? (Myself, so far I do not use Perl's
digested data, and parse Unicode Consortium files directly - so I
do not qualify to judge.) Put the stuff into Unicode::UCD::age?
BTW, why Unicode::UCD has so bizzare interface? Why not have
Unicode::UCD::Name, for example? (The most important piece of data of
those not available via Perl4 interfaces...)
Ilya
P.S. Is unicore/NamesList.txt included with latest distributions of
Perl? My module relies on parsing this file, and... Aha, found it on
http://cpansearch.perl.org/src/FLORA/perl-5.14.2/lib/unicore/
, good!
------------------------------
Date: Tue, 10 Jan 2012 22:02:04 +0000
From: Ben Morrow <ben@morrow.me.uk>
Subject: Re: UCD, and Unicode-AGE of a character?
Message-Id: <s3put8-8vt.ln1@anubis.morrow.me.uk>
Quoth Ilya Zakharevich <nospam-abuse@ilyaz.org>:
> On 2012-01-10, Ben Morrow <ben@morrow.me.uk> wrote:
> > Quoth Ilya Zakharevich <nospam-abuse@ilyaz.org>:
> >> I looked through the docs I could find, and can't find any way to
> >> determine the "Unicode AGE" of a particular codepoint except for:
> >>
> >> a) running /\p{Present_in: FOO}/ for all forseeable values of FOO;
> >>
> >> b) manually parsing $out = do 'unicore/To/Age.pl';.
> >>
> >> Do I miss anything?
> >
> > I don't think so. Note that before (I think) 5.14 unicore/To/Age.pl
> > doesn't exist, and before (I think) 5.12 unicode/DAge.txt doesn't exist
> > either. You may be better off just grabbing a copy of DerivedAge.txt
> > from the Unicode Consortium directly, and using that.
>
> What would be the best fix? (Myself, so far I do not use Perl's
> digested data, and parse Unicode Consortium files directly - so I
> do not qualify to judge.) Put the stuff into Unicode::UCD::age?
For you, or for perl? For your purposes I'd've thought the best thing to
do is ship DerivedAge.txt from Unicode 6.0.0 with your application, then
parse unicore/DAge.txt if it's there and fall back to the shipped copy
if not. That way you've got data for at least the version of Unicode
this perl supports.
For perl, yes, I would have thought Unicode::UCD is the right place, but
I don't really do any serious Unicode work so my opinion doesn't count
for much :).
> BTW, why Unicode::UCD has so bizzare interface? Why not have
> Unicode::UCD::Name, for example?
I don't know, you'd have to ask Jarkko... I certainly agree there's a
lot missing. I suspect that the main effort for 5.8 was to get enough
implemented to make the regex engine work, and since then things have
only been added when someone has been sufficiently motivated to send in
a patch.
> (The most important piece of data of those not available via Perl4
> interfaces...)
Well, they're not strictly 'Perl4' interfaces, of course, since none of
this existed before 5.8...
Ben
------------------------------
Date: Tue, 10 Jan 2012 06:47:56 +0000 (UTC)
From: Ilya Zakharevich <nospam-abuse@ilyaz.org>
Subject: Unicode-AGE of a character?
Message-Id: <slrnjgnnos.4ik.nospam-abuse@panda.math.berkeley.edu>
I looked through the docs I could find, and can't find any way to
determine the "Unicode AGE" of a particular codepoint except for:
a) running /\p{Present_in: FOO}/ for all forseeable values of FOO;
b) manually parsing $out = do 'unicore/To/Age.pl';.
Do I miss anything?
Thanks,
Ilya
------------------------------
Date: Tue, 10 Jan 2012 12:02:43 +0000
From: Ben Morrow <ben@morrow.me.uk>
Subject: Re: Unicode-AGE of a character?
Message-Id: <30mtt8-som.ln1@anubis.morrow.me.uk>
Quoth Ilya Zakharevich <nospam-abuse@ilyaz.org>:
> I looked through the docs I could find, and can't find any way to
> determine the "Unicode AGE" of a particular codepoint except for:
>
> a) running /\p{Present_in: FOO}/ for all forseeable values of FOO;
>
> b) manually parsing $out = do 'unicore/To/Age.pl';.
>
> Do I miss anything?
I don't think so. Note that before (I think) 5.14 unicore/To/Age.pl
doesn't exist, and before (I think) 5.12 unicode/DAge.txt doesn't exist
either. You may be better off just grabbing a copy of DerivedAge.txt
from the Unicode Consortium directly, and using that.
Ben
------------------------------
Date: 6 Apr 2001 21:33:47 GMT (Last modified)
From: Perl-Users-Request@ruby.oce.orst.edu (Perl-Users-Digest Admin)
Subject: Digest Administrivia (Last modified: 6 Apr 01)
Message-Id: <null>
Administrivia:
To submit articles to comp.lang.perl.announce, send your article to
clpa@perl.com.
Back issues are available via anonymous ftp from
ftp://cil-www.oce.orst.edu/pub/perl/old-digests.
#For other requests pertaining to the digest, send mail to
#perl-users-request@ruby.oce.orst.edu. Do not waste your time or mine
#sending perl questions to the -request address, I don't have time to
#answer them even if I did know the answer.
------------------------------
End of Perl-Users Digest V11 Issue 3586
***************************************