[25896] in Perl-Users-Digest

home help back first fref pref prev next nref lref last post

Perl-Users Digest, Issue: 8124 Volume: 10

daemon@ATHENA.MIT.EDU (Perl-Users Digest)
Sat May 28 11:05:35 2005

Date: Sat, 28 May 2005 08:05:06 -0700 (PDT)
From: Perl-Users Digest <Perl-Users-Request@ruby.OCE.ORST.EDU>
To: Perl-Users@ruby.OCE.ORST.EDU (Perl-Users Digest)

Perl-Users Digest           Sat, 28 May 2005     Volume: 10 Number: 8124

Today's topics:
        How to clear all html tag in document? <max@xxxxxxx.xxx>
    Re: How to clear all html tag in document? <jurgenex@hotmail.com>
    Re: How to clear all html tag in document? <nobull@mail.com>
    Re: new how-to book about tit-fucking <wallace259@hotmail.com>
    Re: output question <vdotniekerkatfreelerdotnl>
    Re: Scanning @array elements for similair content <rbutcher.nospam@hotmail.com>
    Re: Scanning @array elements for similair content <jurgenex@hotmail.com>
    Re: Scanning @array elements for similair content <rbutcher.nospam@hotmail.com>
    Re: Scanning @array elements for similair content <djames@thehub.com.au>
    Re: Scanning @array elements for similair content <nobull@mail.com>
        Digest Administrivia (Last modified: 6 Apr 01) (Perl-Users-Digest Admin)

----------------------------------------------------------------------

Date: Sat, 28 May 2005 08:48:27 +0200
From: "max" <max@xxxxxxx.xxx>
Subject: How to clear all html tag in document?
Message-Id: <d7947p$c43$1@ss405.t-com.hr>

How to clear all html tag in line (or document) ?
All tags have "<" on start, and ">" at the end of tag. Eg <table>, </table>,
<div align="left">, <td align="left" bgcolor="#000099"> ....
I make program that work character by character, and control if is start is
"<", and end ">".
Please help me I now those Perl programmers do that on easier way! How?

Thanks




------------------------------

Date: Sat, 28 May 2005 07:15:26 GMT
From: "Jürgen Exner" <jurgenex@hotmail.com>
Subject: Re: How to clear all html tag in document?
Message-Id: <iUUle.2468$vK5.576@trnddc03>

max wrote:
> How to clear all html tag in line (or document) ?

Is there anything wrong with the answer in the FAQ  'perldoc -q "remove 
HTML"'
    "How do I remove HTML from a string?"

> Please help me I now those Perl programmers do that on easier way!
> How?

Trivial. They just follow the suggestions in the FAQ.

jue 




------------------------------

Date: Sat, 28 May 2005 08:16:47 +0100
From: Brian McCauley <nobull@mail.com>
Subject: Re: How to clear all html tag in document?
Message-Id: <d795su$cgh$2@slavica.ukpost.com>



max wrote:

> How to clear all html tag in line (or document) ?

See FAQ.


------------------------------

Date: Sat, 28 May 2005 11:49:19 GMT
From: "LarryW" <wallace259@hotmail.com>
Subject: Re: new how-to book about tit-fucking
Message-Id: <3VYle.2304$p94.1074@tornado.ohiordc.rr.com>

IkFnZW50IDY5IiA8Njktbm8tc3BhbUA2OS42OS42OS42OS5pbnZhbGlkPiB3cm90ZSBpbiBtZXNz
YWdlIG5ld3M6Y0RJek1UQT0uY2Y1ZmMwYWZjOTcxMzFhYzYyMjU5ZTlhZWY1NDhkN2JAMTExNzIy
MzQ0Ni5udWxsdXNlci5jb20uLi4NCj4gDQo+IC0tIA0KPiBBZ2VudCA2OQ0KPiBhbHQuZHJ1Z3Mu
aHlkcm9tb3JwaG9uZQ0KPiANCj4gIi4uLnNvbWUgc2F5IHRydWUgY29tZWR5IGlzIHRoZSB3b3Jr
IG9mIGdlbml1cy4gSSBkaXNhZ3JlZSwgdGhlIA0KPiBmdW5uaWVzdCBzaGl0IEkgZXZlciBzYXcg
d2FzIGFsbCB0aGUgd29yayBvZiBoYWxmd2l0cy4iDQo+IC0tIFNpLCBpbiBhbHQuZHJ1Z3MuaGFy
ZCwgNS4gTWF5IDIwMDUNCj4gDQo+DQpTZWVraW5nIGdvb2QgbWVkaWEgYnJvYWRjYXN0IGNvbWVk
eT8gRm9yZ2V0IGFsbCB0aGUgc2l0Y29tcywgbW9ub2xvZ3VlcyBhbmQgb3RoZXIgY2hlYXAgc2Ny
aXB0ZWQgbWF0ZXJpYWwuIEknbGwgcmF0aGVyIGEgZ29vZCBkb3NlIG9mIDxzZWxlY3Q+IHJlYWxp
dHkgVFYuIFNwcmluZ2VyIHdhcyB0aGUgcGlvbmVlciBpbiB0aGUgZGlzY292ZXJ5IG9mICJ3b3Jr
IG9mIGhhbGZ3aXRzLiIgVGhlIHNob3cgIkNvcHMiIGFuZCAiSnVkZ2UgSnVkeSIgYXJlIHR3byBl
eGFtcGxlcyBvZiB0cnVlIGdlbXMsIGJ1dCB5b3UndmUgZ290dGEgbm90IGJlIGRpc3RyYWN0ZWQg
YnkgdGhlIGluaGVyZW50IHByZWRpY3RhYmxlIGRyYW1hIHRvIGJlIHNoYXJwIGVub3VnaCB0byBw
aWNrIHVwIG9uIHRoZSBmbG9vZCBvZiA8Z2VudWluZT4gaW1wcm9tcHR1IGNvbWVkeSBvZiBsb2dp
Yy4gVGhlIHNjYXJ5IHBhcnQgb2YgaXQgaXMgdGhlIGJ1bGsgb2YgdGhlIFRWIHdhdGNoZXJzIGhh
dmUgdGhlIGxpbWl0ZWQgY2FwYWNpdHkgdG8gc2VlIG9ubHkgdGhlIGRyYW1hLiBIYWxmd2l0cyBz
aW1wbHkgY2FuJ3QgbGF1Z2ggYXQgdGhlbXNlbHZlcy4NCg0KTGFycnlXDQotLSANCg0KDQoiVGhl
IG9ubHkgY2VydGFpbnR5IGlzIHRoYXQgYWxsIHRoaW5ncyB3aWxsIHBhc3MuIg==



------------------------------

Date: Sat, 28 May 2005 11:12:42 +0200
From: Huub <vdotniekerkatfreelerdotnl>
Subject: Re: output question
Message-Id: <429835de$0$151$3a628fcd@reader2.nntp.hccnet.nl>

Tad McClellan wrote:
> Huub <> wrote:
> 
> 
>>>You do know that you don't need format() to make output, don't you?
>>
>> From the tutorial I'm studying, I understood I needed format() 
>>to write to an outputfile.
> 
> 
> 
> Sounds like you are attempting to learn from a foolish tutorial
> (there are lots of those).
> 
> Which tutorial are you studying?
> 
> 

Introduction to Perl & CGI Programming
A complete course of study at free−ed.net
Copyright © 1999, Free−Ed, Ltd.
All rights reserved.


------------------------------

Date: Sat, 28 May 2005 05:47:37 GMT
From: "Randy" <rbutcher.nospam@hotmail.com>
Subject: Re: Scanning @array elements for similair content
Message-Id: <ZBTle.1500813$6l.598100@pd7tw2no>

"Tad McClellan" <tadmc@augustmail.com> wrote in message
news:slrnd9fku9.e81.tadmc@magna.augustmail.com...
> Randy <rbutcher.nospam@hotmail.com> wrote:

>    perldoc -q duplicate
>
>        How can I remove duplicate elements from a list or array?
>
> (pay particular attention to the last sentence of the answer given there.)
>
> You are expected to check the Perl FAQ *before* posting to
> the Perl newsgroup you know.


I had actually checked the perldoc for this and did find ways to remove
duplicate array entires but didn't know what to do when I wanted to match
the specific split item .. ie .. match the email address only, not the
entire array element.


> > use strict;

> Very good, but you should also have:
>
>    use warnings;
>
> (and look in your server error logs for its output, or,
>  even better, run your CGI program from the command line
>  during early development, rather than in the CGI environment)


I have also added 'use warnings,' thank you.


> > open (FH, "<data.txt") or die "Can't open file: $!";
> >   @data=<FH>;
> > close(FH);
> >
> > foreach $data (@data) {
>
>
> It is bad practice to read an entire file into memory only to
> process it line-by-line anyway.
>
> Why not just read and process line-by-line?


Agreed, I now use the CPAN module File::Slurp to read textfile entires into
an @array efficiently.
http://search.cpan.org/~uri/File-Slurp-9999.09/lib/File/Slurp.pm


> Whitespace is not a scarce resource, feel free to use as much of
> it as you like to make your code easier to read.


Point noted.


> > Another complication is
> > what if there are two identical email addresses but one is all caps and
the
> > other isn't.
>
> You need to decide what to do, then we can help you write Perl
> code that does that.
>
> You could perhaps just normalize them all to a single case
> before storing or searching the hash:


During the split phase where I separate the $name for the $email, I now use
this regex: $email =~ tr/A-Z/a-z/;


> > instead to point me in the right direction so that I actually learn
> > something and forward my Perl skills.
>
> A depressingly infrequent display of Good Attitude for this here group.
>
> Good for you!  (and us)


You are correct Tad, I am new to programming in general. I'm trying my best
to better understand the basics. Here is the final code I use to remove the
duplicate entries and it does do it's job:

#!/usr/bin/perl

use CGI;
use CGI::Carp qw(fatalsToBrowser);

use strict;
use warnings;

open (INPUT, "<data.txt") or die "Can't open file: $!";
  my %entries;
  while ( my $data = <INPUT> ) {
    chomp $data;
    my ($name, $email) = split(/\,/, $data);
    $name =~ s/(\w+)/\u\L$1/g;
    $email =~ tr/A-Z/a-z/;
    $entries{$email} = $name;
  }
close(INPUT);

foreach my $adr ( sort keys %entries ) {
  print "$entries{$adr},$adr\n";
}

exit;

That said, I'm not entirely certain what part of the code IS actually
detecting and removing the duplicate entries. I have a hunch that this is
taking please in the foreach loop. I created a test data.txt file and
manually entered several duplicate email addresses. When the script is run,
any duplicate is removed, seems it kills duplicates from the top down .. ie
 .. if dan@email.com was found on 5 lines, it keeps the last occurrence... or
maybe it removed all but the last alphabetically sorted item.

Tad, thank you for this. I would like to ask one final question on this
matter ... right now, when the script is run, it prints to screen all remain
hash entries without any duplicates. Under that I would like it to show
which entries got removed. I assume to do this, I would need to modify the
script to push any matched duplicates into a secondary array or hash and
then print that last. Perhaps not. Your thoughts are appreciated.

Robert

P.S. Your going to laugh at this but until recent I have never used the
command 'use strict'. To be honest, I'm not 100% certain what exactly this
does, or how it is benefiting me or the script. All I know for certain is
that without adding "my" to variable definitions, the script doesn't
work/run. Most articles I have read online highly recommend using this
command but don't go into great detail why. I ask you this because I wish to
better my understanding of Perl and to ensure I write proper scripts in the
future.




------------------------------

Date: Sat, 28 May 2005 06:13:44 GMT
From: "Jürgen Exner" <jurgenex@hotmail.com>
Subject: Re: Scanning @array elements for similair content
Message-Id: <s_Tle.5545$3u3.4481@trnddc07>

Randy wrote:
> During the split phase where I separate the $name for the $email, I
> now use this regex: $email =~ tr/A-Z/a-z/;

Just to be nitpicking: tr/// does not use REs. That's one big difference to 
s///.

And it's better to use the function lc() instead of your tr/// code because 
lc() handles non-English characters correctly, too, while your code fails 
for anything that is outside the basic 26 latin characters.

jue 




------------------------------

Date: Sat, 28 May 2005 06:36:15 GMT
From: "Randy" <rbutcher.nospam@hotmail.com>
Subject: Re: Scanning @array elements for similair content
Message-Id: <zjUle.1497823$Xk.954417@pd7tw3no>

"Randy" <rbutcher.nospam@hotmail.com> wrote in message
news:ZBTle.1500813$6l.598100@pd7tw2no...
> "Tad McClellan" <tadmc@augustmail.com> wrote in message

> That said, I'm not entirely certain what part of the code IS actually
> detecting and removing the duplicate entries. I have a hunch that this is
> taking please in the foreach loop.

Tad, I did a little more research on hashes. I now think the duplicate
elimination is NOT happening during the foreach loop, that loop is just
sorting the hash and printing it; instead it is occuring when you are
defining each hash element in the initial <while> loop. I think this happens
because you are assigning (your method) the key as the email address and the
value as the name. In doing so you can't have duplicated key names?!?!?! so
the hash just ignores when a request for a duplicate key name is
requested?!?!?!?

If i'm wrong about this I hope you don't think less of me ... I really am
trying to learn.

Robert




------------------------------

Date: Sat, 28 May 2005 06:46:28 GMT
From: Damian James <djames@thehub.com.au>
Subject: Re: Scanning @array elements for similair content
Message-Id: <slrnd9g4u2.eki.djames@puli.local>

On Sat, 28 May 2005 06:36:15 GMT, Randy said:
> ... so
> the hash just ignores when a request for a duplicate key name is
> requested?!?!?!?

No, the second assignment simply overrides the first one.

my %blah;
$blah{ x } = 'test 1';
$blah{ x } = 'test 2';
print "$blah{ x }\n";

> If i'm wrong about this I hope you don't think less of me ... I really am
> trying to learn.

Pfft, we've all been there. Never care about seeming foolish when
the object is to learn. It's the folks who try to look like they
already know everything who are foolish.

--damian


------------------------------

Date: Sat, 28 May 2005 08:09:59 +0100
From: Brian McCauley <nobull@mail.com>
Subject: Re: Scanning @array elements for similair content
Message-Id: <d795gk$cgh$1@slavica.ukpost.com>



Randy wrote:

> "Tad McClellan" <tadmc@augustmail.com> wrote in message
> news:slrnd9fku9.e81.tadmc@magna.augustmail.com...
> 
> 
>>   perldoc -q duplicate
>>
>>       How can I remove duplicate elements from a list or array?
>>
>>(pay particular attention to the last sentence of the answer given there.)
>>
>>You are expected to check the Perl FAQ *before* posting to
>>the Perl newsgroup you know.
> 
> I had actually checked the perldoc for this and did find ways to remove
> duplicate array entires but didn't know what to do when I wanted to match
> the specific split item .. ie .. match the email address only, not the
> entire array element.

 >>Randy <rbutcher.nospam@hotmail.com> wrote:
 >>
>>>open (FH, "<data.txt") or die "Can't open file: $!";
>>>  @data=<FH>;
>>>close(FH);
>>>
>>>foreach $data (@data) {
>>
>>It is bad practice to read an entire file into memory only to
>>process it line-by-line anyway.
>>
>>Why not just read and process line-by-line?
> 
> Agreed, I now use the CPAN module File::Slurp to read textfile entires into
> an @array efficiently.

If you have a need to slurp then File::Slurp will do so efficiently but 
you have no need.  It is better to read a line at a time as Tad showed. 
I see looking a the end of the post you have indeed done so.  Good.

>>> Another complication is what if there are two identical
 >>> email addresses but one is all caps and
>>> the other isn't.
>>
>>You could perhaps just normalize them all to a single case
>>before storing or searching the hash:
> 
> During the split phase where I separate the $name for the $email, I now use
> this regex: $email =~ tr/A-Z/a-z/;

There is no regex there.  I agree with Tad that the lc() function would 
be beter than tr///.

     $email = lc $email;

>>>instead to point me in the right direction so that I actually learn
>>>something and forward my Perl skills.
>>
>>Good for you!  (and us)

Ditto.

> You are correct Tad, I am new to programming in general. I'm trying my best
> to better understand the basics. Here is the final code I use to remove the
> duplicate entries and it does do it's job:

It looks good.  I will now proceed t criticise it but don't let this 
detract from the fact that it is good.

> #!/usr/bin/perl
> 
> use CGI;
> use CGI::Carp qw(fatalsToBrowser);

Is this a CGI script?  It doesn't look time one?

> use strict;
> use warnings;

Generally best to put these two ASAP.  That way you'll even get their 
protection in your other use statements.  The only thing I like to see 
above these two are comments and, in a the case of a module, a package 
directive.

> open (INPUT, "<data.txt") or die "Can't open file: $!";
>   my %entries;
>   while ( my $data = <INPUT> ) {
>     chomp $data;
>     my ($name, $email) = split(/\,/, $data);

No need to backslash the comma in a regex.  I'm not as paranoid about 
leaning toothpick syndrome as Tad but I wouldn't bother here.

>     $name =~ s/(\w+)/\u\L$1/g;

OK, nothing whatever to do with Perl, but this is bad.  There are a lot 
of names (like mine) that have non-trivial capitaliztion.  You risk 
offending and alienating many people.  This has been oft discssed here. 
  There is no solution as sometimes there can be two distinct names that 
differ only in capialization.

>     $email =~ tr/A-Z/a-z/;
>     $entries{$email} = $name;
>   }
> close(INPUT);

Your code looks nice but your use of indentation between the open/close 
is rather unconventional.

> foreach my $adr ( sort keys %entries ) {
>   print "$entries{$adr},$adr\n";
> }
> 
> exit;

It is more conventional just to let perl fall off the end of your script 
and exit() implicitly.

> That said, I'm not entirely certain what part of the code IS actually
> detecting and removing the duplicate entries. I have a hunch that this is
> taking please in the foreach loop.

No - it is the line

      $entries{$email} = $name;

If you encounter a second record in the input with an e-mail address 
that's been encountered before the above line will replace the old entry 
in %entries with a new one, thus forgetting all but the last entry with 
a given e-mail.

> .. if dan@email.com was found on 5 lines, it keeps the last occurrence...

Yep.

> Tad, thank you for this. I would like to ask one final question on this
> matter ... right now, when the script is run, it prints to screen all remain
> hash entries without any duplicates. Under that I would like it to show
> which entries got removed. I assume to do this, I would need to modify the
> script to push any matched duplicates into a secondary array or hash and
> then print that last.

Yes that would work.

   if ( defined $entries{$email} ) {
     push @duplicates => $data;
   } else {
     $entries{$email} = $name;
   }

Note - this now preserves the first instance of each address and puts 
the rest into @duplicates.

> P.S. Your going to laugh at this but until recent I have never used the
> command 'use strict'. To be honest, I'm not 100% certain what exactly this
> does, or how it is benefiting me or the script. All I know for certain is
> that without adding "my" to variable definitions, the script doesn't
> work/run.

Yes that is probably the most noticable of the three effects. Without 
'use strict' perl will treat the first mention of an undeclared variable 
as an implicit declaration of a package-scoped variable (well kinda). 
This can be a great convenience in 1-line scripts but is generally a 
liability in scripts longer than about 10 lines.

> Most articles I have read online highly recommend using this
> command but don't go into great detail why. I ask you this because I wish to
> better my understanding of Perl and to ensure I write proper scripts in the
> future.

I would argue (and indeed have argued with giants) that it is best to 
see 'use strict' as disabling three fairly obscure features and that 
understanding of these features is something that should not concern 
people too early in their learning of Perl.

http://groups-beta.google.com/group/comp.lang.perl.misc/msg/89f307d6b9e83c65


------------------------------

Date: 6 Apr 2001 21:33:47 GMT (Last modified)
From: Perl-Users-Request@ruby.oce.orst.edu (Perl-Users-Digest Admin) 
Subject: Digest Administrivia (Last modified: 6 Apr 01)
Message-Id: <null>


Administrivia:

#The Perl-Users Digest is a retransmission of the USENET newsgroup
#comp.lang.perl.misc.  For subscription or unsubscription requests, send
#the single line:
#
#	subscribe perl-users
#or:
#	unsubscribe perl-users
#
#to almanac@ruby.oce.orst.edu.  

NOTE: due to the current flood of worm email banging on ruby, the smtp
server on ruby has been shut off until further notice. 

To submit articles to comp.lang.perl.announce, send your article to
clpa@perl.com.

#To request back copies (available for a week or so), send your request
#to almanac@ruby.oce.orst.edu with the command "send perl-users x.y",
#where x is the volume number and y is the issue number.

#For other requests pertaining to the digest, send mail to
#perl-users-request@ruby.oce.orst.edu. Do not waste your time or mine
#sending perl questions to the -request address, I don't have time to
#answer them even if I did know the answer.


------------------------------
End of Perl-Users Digest V10 Issue 8124
***************************************


home help back first fref pref prev next nref lref last post